Alert when an integration or DNS provider fails repeatedly

The Diagnostic Log already records every outbound call's success or
failure, but nothing acted on it — you'd only notice an integration
was down by happening to open its page. Adds a per-source consecutive-
failure counter (in-memory, reset on restart, same durability tier as
the diag log's own ring buffer) hooked into recordDiagEntry: crossing
the configurable threshold (default 3) sends one "down" notification
on every configured channel, and a "recovered" notification fires once
it succeeds again — no repeat spam while it stays down. New
"Integration/DNS provider failing repeatedly" toggle and threshold
field under Settings -> Notifications.

Verified end-to-end against an isolated scratch database with a real
local HTTP server standing in for the webhook channel: 5 consecutive
failures produced exactly one "Down" notification (at the 3rd
failure, correctly naming "3 calls"), a subsequent success produced
exactly one "Recovered" notification, and two more failures on a
fresh streak triggered nothing (below threshold) — confirmed the real
dev database's mtime was untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
bobbanandClaude Sonnet 5 committed 2026-09-19 14:51:00 +02:00
1 parent 10d123b18a
commit b5a4c6e2d9
7 files changed
+105

No files matched your search

@@ -0,0 +1,38 @@
import { getSettings } from "./settingsStore.js";
import { notifyIntegrationDown, notifyIntegrationRecovered } from "./notify.js";
interface SourceHealth {
consecutiveFailures: number;
/** True once a "down" notification has been sent for the current failure streak, so it isn't repeated on every subsequent failure. */
alerted: boolean;
}
/**
* In-memory only (like the diag log's own ring buffer — resets on restart,
* which is fine since this tracks a live streak, not history). Keyed by diag
* log "source" (e.g. "proxmox", "cloudflare") — the same granularity the
* Diagnostic Log itself filters by, so a homelab running two integrations of
* the same type shares one health streak between them.
*/
const health = new Map<string, SourceHealth>();
/** Called after every diagnostic-log entry is recorded, to track consecutive failures per source and alert on threshold-cross / recovery. */
export async function trackIntegrationHealth(source: string, ok: boolean): Promise<void> {
const state = health.get(source) ?? { consecutiveFailures: 0, alerted: false };
if (ok) {
if (state.alerted) {
await notifyIntegrationRecovered(source);
}
health.set(source, { consecutiveFailures: 0, alerted: false });
return;
}
state.consecutiveFailures += 1;
const { notifications } = await getSettings();
if (notifications.integrationFailureAlerts && !state.alerted && state.consecutiveFailures >= notifications.integrationFailureThreshold) {
state.alerted = true;
await notifyIntegrationDown(source, state.consecutiveFailures);
}
health.set(source, state);
}