Add maintenance mode to silence alerts while working on a server or integration
Rebooting Proxmox or patching a server triggered failure/offline alerts you then had to dismiss. A maintenance window silences alerts about one server, integration, or DNS provider for a chosen time. New Maintenance page (start with a duration and optional reason, end early, see what's silenced and what isn't) and a banner in the app shell so every signed-in user can see what is currently silenced. Starting/ending is operator-only and audit-logged; starting one on a target that already has a window restarts its clock instead of stacking. Silenced for the target: server offline/disk alerts, Proxmox/Synology storage and health alerts, Proxmox backup alerts, and "integration down" alerts. Not silenced: expiry and update reminders, DNS change notices. The design goal is that this cannot hide a real outage: - Every window has a required end (5 min to 7 days); there is no open-ended option, so a forgotten window expires by itself. - A silenced problem is deliberately NOT recorded as "known". If it is still present when the window ends it alerts then, as new. A problem that was already alerted before the window stays known, so it isn't repeated, and is reported cleared only after the window ends. - Failure alerts keep counting failures during a window without marking themselves alerted, so an outage that outlasts the window alerts on the very next failed call. Known limitation, stated on the page: integration-failure alerts are tracked per service TYPE (all "proxmox"), not per configured instance, so a window on one Proxmox integration also silences a failure on a second Proxmox integration while it's open. Fixing that means threading the integration id through every adapter and the diagnostic log, which is a much larger change than this feature. Also moved the API-error-message helper out of Secrets.tsx into a shared util now that two pages use it. New table maintenance_windows (migration 0008). Verified with 44 checks: the condition-key-to-subject mapping (including server:3 vs server:33), the diff rules with silenced subjects (new problem not recorded, alerts when the window ends; already-known one carried and not repeated; clears only after the window), window expiry and integration/DNS-provider source matching, the failure tracker end to end against a webhook (silent during a window while an unrelated service still alerts; outage that outlasts the window alerts on the next failure and only once; fail-and-recover fully inside a window sends nothing), a full health pass against a real window, and the real router with a stubbed session (role rules, duration bounds including the missing-duration case, extend-not-stack, 404s, deleted targets hidden, audit entries). Real dev database mtime untouched. Not done: I haven't clicked through the new page or banner in a browser (they sit behind the Authentik login); it builds and the API behind it is tested. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
1 parent
1688de3ea2
commit
aae4f0d74f
18 files changed
+1854
-26
No files matched your search
@@ -1,4 +1,5 @@
|
||||
import { getSettings } from "./settingsStore.js";
|
||||
import { isSourceInMaintenance } from "./maintenance.js";
|
||||
import { notifyIntegrationDown, notifyIntegrationRecovered } from "./notify.js";
|
||||
|
||||
interface SourceHealth {
|
||||
@@ -31,8 +32,12 @@ export async function trackIntegrationHealth(source: string, ok: boolean): Promi
|
||||
state.consecutiveFailures += 1;
|
||||
const { notifications } = await getSettings();
|
||||
if (notifications.integrationFailureAlerts && !state.alerted && state.consecutiveFailures >= notifications.integrationFailureThreshold) {
|
||||
state.alerted = true;
|
||||
await notifyIntegrationDown(source, state.consecutiveFailures);
|
||||
// During maintenance the failures keep being counted but `alerted` stays false, so if the service is
|
||||
// still failing once the window ends, the very next failed call alerts — a real outage isn't swallowed.
|
||||
if (!(await isSourceInMaintenance(source))) {
|
||||
state.alerted = true;
|
||||
await notifyIntegrationDown(source, state.consecutiveFailures);
|
||||
}
|
||||
}
|
||||
health.set(source, state);
|
||||
}
|
||||
Reference in new issue
Block a user