Alert when an integration or DNS provider fails repeatedly

The Diagnostic Log already records every outbound call's success or
failure, but nothing acted on it — you'd only notice an integration
was down by happening to open its page. Adds a per-source consecutive-
failure counter (in-memory, reset on restart, same durability tier as
the diag log's own ring buffer) hooked into recordDiagEntry: crossing
the configurable threshold (default 3) sends one "down" notification
on every configured channel, and a "recovered" notification fires once
it succeeds again — no repeat spam while it stays down. New
"Integration/DNS provider failing repeatedly" toggle and threshold
field under Settings -> Notifications.

Verified end-to-end against an isolated scratch database with a real
local HTTP server standing in for the webhook channel: 5 consecutive
failures produced exactly one "Down" notification (at the 3rd
failure, correctly naming "3 calls"), a subsequent success produced
exactly one "Recovered" notification, and two more failures on a
fresh streak triggered nothing (below threshold) — confirmed the real
dev database's mtime was untouched throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
bobbanandClaude Sonnet 5 committed 2026-09-19 14:51:00 +02:00
1 parent 10d123b18a
commit b5a4c6e2d9
7 files changed
+105

No files matched your search

+2
View File
@@ -63,6 +63,8 @@ const updateSchema = z.object({
tailscaleKeyCheck: z.boolean(),
secretCheckTime: z.string().regex(/^\d{2}:\d{2}$/),
timezone: z.string(),
integrationFailureAlerts: z.boolean(),
integrationFailureThreshold: z.number().int().min(1).max(20),
})
.partial()
.optional(),