Alert when a Semaphore template or Gitea workflow run fails

Every 15 minutes (on the existing health-check timer) the app reads the
latest run of each Semaphore template and each Gitea repo's latest
workflow run. A failed one raises one notification, and another when a
later run succeeds. It is state-based like the server health alerts, so a
job that fails every night alerts on the first failure, not every night.
Gitea alerts include the run's link. Toggle: Settings > Notifications.

What counts:
- Semaphore "error" is a failure, "success" is a pass. A run that is
  waiting, running, stopped by hand or rejected is neither, so it leaves
  the previous state alone: a run in progress must not clear a failure it
  hasn't fixed yet, and a manual stop isn't a failure.
- Gitea failure/success likewise; running, waiting, blocked, cancelled and
  skipped leave things as they were.
- A failing template that gets another failing run does not re-alert.

Not mistaking "couldn't read" for "fixed":
- Semaphore's template listing swallowed per-project errors, so a project
  that failed to load looked like a project with no templates. A new
  checkTemplates adapter method reports which projects failed, and their
  failures are held rather than cleared.
- Gitea reports a run it couldn't fetch as null, the same as "no runs";
  both leave the repo's state alone.
- An unreachable integration holds all of its failures. Nothing is cleared
  or re-announced while it is down.

The first pass only records what is already failing without announcing
it, so upgrading (or adding an integration to a fresh install) doesn't
produce a wall of alerts about months-old failures. That baseline is not
spent while nothing could be read.

Maintenance windows on a Semaphore or Gitea integration silence its
failure alerts with the same rules as the health alerts: a problem that
starts during a window alerts when it ends, and one already announced
stays known. The diff logic is reused from the health monitor rather than
copied. Maintenance page text updated.

Known limit: for Gitea this follows the repo's most recent run on any
workflow or branch, matching what the Gitea page shows; a failure in one
workflow can be masked by a later success of another.

Verified with 53 checks against fake Semaphore and Gitea servers and a
webhook receiver: classification, baseline (including not being consumed
when nothing is readable), single alert per failure, no repeat, in-progress/
stopped/cancelled runs, recovery and re-failure, unreadable project,
unreadable integration, run-fetch errors, maintenance windows (silenced,
then announced after), the toggle, disabled integrations and repos without
Actions. Real dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
bobbanandClaude Sonnet 5 committed 2026-09-26 03:11:59 +02:00
1 parent 9f1609c4ed
commit 4118062405
11 files changed
+219 -5

No files matched your search

+11 -3
View File
@@ -48,6 +48,8 @@ export interface SemaphoreTemplate {
export interface SemaphoreAdapter {
ping(): Promise<{ ok: boolean; latencyMs?: number; error?: string }>;
listTemplatesWithStatus(): Promise<SemaphoreTemplate[]>;
/** Like listTemplatesWithStatus, but says which projects couldn't be read — for callers that must not mistake "couldn't read" for "no templates". */
checkTemplates(): Promise<{ templates: SemaphoreTemplate[]; failedProjectIds: number[] }>;
runTemplate(projectId: number, templateId: number): Promise<SemaphoreTask>;
}
@@ -112,18 +114,24 @@ export function createSemaphoreAdapter(config: SemaphoreConfig): SemaphoreAdapte
}));
}
async function listTemplatesWithStatus(): Promise<SemaphoreTemplate[]> {
async function checkTemplates(): Promise<{ templates: SemaphoreTemplate[]; failedProjectIds: number[] }> {
const projects = await listProjects();
const failedProjectIds: number[] = [];
const perProject = await Promise.all(
projects.map(async (p) => {
try {
return await listTemplatesForProject(p.id, p.name);
} catch {
failedProjectIds.push(p.id);
return [];
}
}),
);
return perProject.flat();
return { templates: perProject.flat(), failedProjectIds };
}
async function listTemplatesWithStatus(): Promise<SemaphoreTemplate[]> {
return (await checkTemplates()).templates;
}
async function runTemplate(projectId: number, templateId: number): Promise<SemaphoreTask> {
@@ -141,5 +149,5 @@ export function createSemaphoreAdapter(config: SemaphoreConfig): SemaphoreAdapte
}
}
return withDiagLogging("semaphore", { ping, listTemplatesWithStatus, runTemplate });
return withDiagLogging("semaphore", { ping, listTemplatesWithStatus, checkTemplates, runTemplate });
}
+1
View File
@@ -69,6 +69,7 @@ const updateSchema = z.object({
dockerUpdateCheck: z.boolean(),
proxmoxBackupCheck: z.boolean(),
healthAlerts: z.boolean(),
automationAlerts: z.boolean(),
secretCheckTime: z.string().regex(/^\d{2}:\d{2}$/),
timezone: z.string(),
integrationFailureAlerts: z.boolean(),
+158
View File
@@ -0,0 +1,158 @@
import { and, eq, inArray } from "drizzle-orm";
import { db } from "../db/client.js";
import { integrations } from "../db/schema.js";
import { loadIntegrationConfig } from "../integrations/loadIntegration.js";
import { createSemaphoreAdapter } from "../integrations/semaphore/adapter.js";
import { createGiteaAdapter } from "../integrations/gitea/adapter.js";
import { diffConditions, type ActiveState, type HealthCondition } from "./healthMonitor.js";
import { activeSubjects } from "./maintenance.js";
import { notifyAutomationFailed, notifyAutomationRecovered } from "./notify.js";
import { getInternalFlag, setInternalFlag } from "./settingsStore.js";
const STATE_FLAG = "automationActiveConditions";
/**
* What a piece of automation's most recent run tells us. "unknown" covers everything that is neither a clear
* pass nor a clear failure — still running, waiting, cancelled, skipped, stopped by hand — and means "don't
* change what we were saying": a run in progress must not clear a failure it hasn't yet fixed, and a manual
* stop isn't a failure.
*/
export type RunOutcome = "failed" | "passed" | "unknown";
export interface AutomationItem {
/** Stable identity: the same template/repo always has the same key. */
key: string;
/** Hierarchical source, "semaphore:3:12:45" — held when the run's outcome isn't known. */
source: string;
outcome: RunOutcome;
/** What to say when it has failed. */
message: string;
}
// ─── Classification (pure) ──────────────────────────────────────────────────
export function classifySemaphoreStatus(status: string | null | undefined): RunOutcome {
if (status === "error") return "failed";
if (status === "success") return "passed";
return "unknown"; // waiting, starting, running, stopping, stopped, rejected, confirmed, ...
}
export function classifyGiteaRun(run: { status: string; conclusion: string | null }): RunOutcome {
if (run.conclusion === "failure" || run.status === "failure") return "failed";
if (run.conclusion === "success" || run.status === "success") return "passed";
return "unknown"; // running, waiting, blocked, cancelled, skipped, ...
}
function endedSuffix(end: string | null): string {
if (!end) return "";
const t = Date.parse(end);
return Number.isFinite(t) ? ` (${new Date(t).toISOString().slice(0, 16).replace("T", " ")} UTC)` : "";
}
/** Splits the items into the failures to report and the sources whose state must be left alone this pass. */
export function evaluateAutomation(items: AutomationItem[]): { conditions: HealthCondition[]; held: Set<string> } {
const conditions: HealthCondition[] = [];
const held = new Set<string>();
for (const item of items) {
if (item.outcome === "failed") conditions.push({ key: item.key, source: item.source, message: item.message });
else if (item.outcome === "unknown") held.add(item.source);
}
return { conditions, held };
}
// ─── Collection (I/O) ───────────────────────────────────────────────────────
async function collect(): Promise<{ items: AutomationItem[]; held: Set<string>; readable: number }> {
const items: AutomationItem[] = [];
const held = new Set<string>();
let readable = 0;
const rows = await db
.select({ id: integrations.id, name: integrations.name, type: integrations.type })
.from(integrations)
.where(and(eq(integrations.enabled, true), inArray(integrations.type, ["semaphore", "gitea"])));
for (const row of rows) {
try {
const loaded = await loadIntegrationConfig(row.id);
if (!loaded) continue;
if (row.type === "semaphore") {
const { templates, failedProjectIds } = await createSemaphoreAdapter(loaded.config as any).checkTemplates();
// A project that couldn't be listed contributes no templates — that's "couldn't read", not "all fixed".
for (const pid of failedProjectIds) held.add(`semaphore:${row.id}:${pid}`);
for (const t of templates) {
if (!t.lastTask) continue;
items.push({
key: `automation:semaphore:${row.id}:${t.projectId}:${t.id}`,
source: `semaphore:${row.id}:${t.projectId}:${t.id}`,
outcome: classifySemaphoreStatus(t.lastTask.status),
message: `${row.name} / ${t.projectName} / ${t.name}: run #${t.lastTask.id} failed${endedSuffix(t.lastTask.end)}`,
});
}
} else {
const repos = await createGiteaAdapter(loaded.config as any).listReposWithStatus();
for (const r of repos) {
if (!r.hasActions) continue;
const source = `gitea:${row.id}:${r.fullName}`;
// The adapter reports a run it couldn't fetch as null, the same as "no runs yet" — either way there's nothing to judge.
if (!r.latestRun) {
held.add(source);
continue;
}
const run = r.latestRun;
items.push({
key: `automation:gitea:${row.id}:${r.fullName}`,
source,
outcome: classifyGiteaRun(run),
message: `${row.name} / ${r.fullName}: "${run.displayTitle || "workflow"}" (run #${run.runNumber}) failed${run.headBranch ? ` on ${run.headBranch}` : ""}${run.htmlUrl ? ` — ${run.htmlUrl}` : ""}`,
});
}
}
readable++;
} catch (err) {
console.error(`[automation] couldn't read ${row.type} integration ${row.id}:`, err instanceof Error ? err.message : err);
held.add(`${row.type}:${row.id}`);
}
}
return { items, held, readable };
}
// ─── The scheduled pass ─────────────────────────────────────────────────────
async function loadState(): Promise<ActiveState | null> {
const raw = await getInternalFlag(STATE_FLAG);
if (raw === null) return null;
try {
return JSON.parse(raw);
} catch {
return {};
}
}
/**
* Reports each template/repo whose latest run failed, once, and again when it succeeds. State-based like the
* health monitor: a template that keeps failing every night alerts on the first failure, not every night.
*
* The very first pass only records what is already failing, without announcing it — on a fresh install (or
* the first time this feature runs) that would otherwise be a wall of alerts about failures that are months old.
*/
export async function runAutomationCheck(now: number = Date.now()): Promise<{ added: number; resolved: number; baseline: boolean }> {
const { items, held: readHeld, readable } = await collect();
const stored = await loadState();
// Nothing could be read at all (or nothing is configured): don't spend the "first pass" on an empty picture.
if (stored === null && readable === 0) return { added: 0, resolved: 0, baseline: false };
const { conditions, held: unknownHeld } = evaluateAutomation(items);
const held = new Set([...readHeld, ...unknownHeld]);
const { added, resolved, next } = diffConditions(stored ?? {}, conditions, held, await activeSubjects(new Date(now)));
await setInternalFlag(STATE_FLAG, JSON.stringify(next));
if (stored === null) return { added: 0, resolved: 0, baseline: true };
// State is tracked even with the alert toggle off (the notify functions check it), so turning it back on isn't a flood.
if (added.length > 0) await notifyAutomationFailed(added);
if (resolved.length > 0) await notifyAutomationRecovered(resolved);
return { added: added.length, resolved: resolved.length, baseline: false };
}
+9 -2
View File
@@ -1,4 +1,5 @@
import { runHealthCheck } from "./healthMonitor.js";
import { runAutomationCheck } from "./automationMonitor.js";
// 15 minutes matches the agent's default report interval, so a check never sees "stale" data that's just waiting for the next report.
const INTERVAL_MS = 15 * 60 * 1000;
@@ -6,17 +7,23 @@ const INTERVAL_MS = 15 * 60 * 1000;
let timer: ReturnType<typeof setInterval> | null = null;
async function pass(): Promise<void> {
// Independent, so a failure in one never skips the other.
try {
await runHealthCheck();
} catch (err) {
console.error("[health] check failed:", err);
}
try {
await runAutomationCheck();
} catch (err) {
console.error("[automation] check failed:", err);
}
}
export async function initHealthScheduler(): Promise<void> {
if (timer) return;
timer = setInterval(() => void pass(), INTERVAL_MS);
// Not awaited: the first pass talks to Proxmox/Synology and must not delay server startup.
// Not awaited: the first pass talks to Proxmox/Synology/Semaphore/Gitea and must not delay server startup.
void pass();
console.log(`[health] checking server/storage health every ${INTERVAL_MS / 60000} minutes`);
console.log(`[health] checking server/storage health and automation runs every ${INTERVAL_MS / 60000} minutes`);
}
+2
View File
@@ -22,6 +22,8 @@ export function subjectOfConditionKey(key: string): string | null {
if (m) return `integration:${m[1]}`;
m = /^(?:synology-volume|synology-disk|disk:synology):(\d+):/.exec(key);
if (m) return `integration:${m[1]}`;
m = /^automation:(?:semaphore|gitea):(\d+):/.exec(key);
if (m) return `integration:${m[1]}`;
return null;
}
+18
View File
@@ -295,6 +295,24 @@ export async function notifyHealthRecovered(resolved: { message: string }[]): Pr
);
}
export async function notifyAutomationFailed(issues: { message: string }[]): Promise<void> {
if (issues.length === 0) return;
if (!(await eventEnabled("automationAlerts"))) return;
await notify(
"Homelab Manager — Automation Failed",
`${issues.length} run${issues.length !== 1 ? "s" : ""} failed:\n\n${issues.map((i) => `✕ ${i.message}`).join("\n")}`,
);
}
export async function notifyAutomationRecovered(resolved: { message: string }[]): Promise<void> {
if (resolved.length === 0) return;
if (!(await eventEnabled("automationAlerts"))) return;
await notify(
"Homelab Manager — Automation Recovered",
`${resolved.length} no longer failing:\n\n${resolved.map((i) => `✓ ${i.message}`).join("\n")}`,
);
}
export async function notifyTailscaleKeyExpiry(
expiring: { integrationName: string; deviceLabel: string; daysLeft: number }[],
): Promise<void> {
+3
View File
@@ -43,6 +43,8 @@ export interface NotificationEvents {
dockerUpdateCheck: boolean;
proxmoxBackupCheck: boolean;
healthAlerts: boolean;
/** A Semaphore template or Gitea repo whose latest run failed, and when it succeeds again. */
automationAlerts: boolean;
secretCheckTime: string; // "HH:MM" — shared by the secret-expiry, Tailscale key-expiry, Docker update, and Proxmox backup checks
timezone: string;
integrationFailureAlerts: boolean;
@@ -114,6 +116,7 @@ const DEFAULTS: AppSettings = {
dockerUpdateCheck: true,
proxmoxBackupCheck: true,
healthAlerts: true,
automationAlerts: true,
secretCheckTime: "08:00",
timezone: "UTC",
integrationFailureAlerts: true,