Alert when an integration or DNS provider fails repeatedly
The Diagnostic Log already records every outbound call's success or failure, but nothing acted on it — you'd only notice an integration was down by happening to open its page. Adds a per-source consecutive- failure counter (in-memory, reset on restart, same durability tier as the diag log's own ring buffer) hooked into recordDiagEntry: crossing the configurable threshold (default 3) sends one "down" notification on every configured channel, and a "recovered" notification fires once it succeeds again — no repeat spam while it stays down. New "Integration/DNS provider failing repeatedly" toggle and threshold field under Settings -> Notifications. Verified end-to-end against an isolated scratch database with a real local HTTP server standing in for the webhook channel: 5 consecutive failures produced exactly one "Down" notification (at the 3rd failure, correctly naming "3 calls"), a subsequent success produced exactly one "Recovered" notification, and two more failures on a fresh streak triggered nothing (below threshold) — confirmed the real dev database's mtime was untouched throughout. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
1 parent
10d123b18a
commit
b5a4c6e2d9
7 files changed
+105
No files matched your search
@@ -0,0 +1,38 @@
|
||||
import { getSettings } from "./settingsStore.js";
|
||||
import { notifyIntegrationDown, notifyIntegrationRecovered } from "./notify.js";
|
||||
|
||||
interface SourceHealth {
|
||||
consecutiveFailures: number;
|
||||
/** True once a "down" notification has been sent for the current failure streak, so it isn't repeated on every subsequent failure. */
|
||||
alerted: boolean;
|
||||
}
|
||||
|
||||
/**
|
||||
* In-memory only (like the diag log's own ring buffer — resets on restart,
|
||||
* which is fine since this tracks a live streak, not history). Keyed by diag
|
||||
* log "source" (e.g. "proxmox", "cloudflare") — the same granularity the
|
||||
* Diagnostic Log itself filters by, so a homelab running two integrations of
|
||||
* the same type shares one health streak between them.
|
||||
*/
|
||||
const health = new Map<string, SourceHealth>();
|
||||
|
||||
/** Called after every diagnostic-log entry is recorded, to track consecutive failures per source and alert on threshold-cross / recovery. */
|
||||
export async function trackIntegrationHealth(source: string, ok: boolean): Promise<void> {
|
||||
const state = health.get(source) ?? { consecutiveFailures: 0, alerted: false };
|
||||
|
||||
if (ok) {
|
||||
if (state.alerted) {
|
||||
await notifyIntegrationRecovered(source);
|
||||
}
|
||||
health.set(source, { consecutiveFailures: 0, alerted: false });
|
||||
return;
|
||||
}
|
||||
|
||||
state.consecutiveFailures += 1;
|
||||
const { notifications } = await getSettings();
|
||||
if (notifications.integrationFailureAlerts && !state.alerted && state.consecutiveFailures >= notifications.integrationFailureThreshold) {
|
||||
state.alerted = true;
|
||||
await notifyIntegrationDown(source, state.consecutiveFailures);
|
||||
}
|
||||
health.set(source, state);
|
||||
}
|
||||
Reference in new issue
Block a user