Alert when a Semaphore template or Gitea workflow run fails

Every 15 minutes (on the existing health-check timer) the app reads the
latest run of each Semaphore template and each Gitea repo's latest
workflow run. A failed one raises one notification, and another when a
later run succeeds. It is state-based like the server health alerts, so a
job that fails every night alerts on the first failure, not every night.
Gitea alerts include the run's link. Toggle: Settings > Notifications.

What counts:
- Semaphore "error" is a failure, "success" is a pass. A run that is
  waiting, running, stopped by hand or rejected is neither, so it leaves
  the previous state alone: a run in progress must not clear a failure it
  hasn't fixed yet, and a manual stop isn't a failure.
- Gitea failure/success likewise; running, waiting, blocked, cancelled and
  skipped leave things as they were.
- A failing template that gets another failing run does not re-alert.

Not mistaking "couldn't read" for "fixed":
- Semaphore's template listing swallowed per-project errors, so a project
  that failed to load looked like a project with no templates. A new
  checkTemplates adapter method reports which projects failed, and their
  failures are held rather than cleared.
- Gitea reports a run it couldn't fetch as null, the same as "no runs";
  both leave the repo's state alone.
- An unreachable integration holds all of its failures. Nothing is cleared
  or re-announced while it is down.

The first pass only records what is already failing without announcing
it, so upgrading (or adding an integration to a fresh install) doesn't
produce a wall of alerts about months-old failures. That baseline is not
spent while nothing could be read.

Maintenance windows on a Semaphore or Gitea integration silence its
failure alerts with the same rules as the health alerts: a problem that
starts during a window alerts when it ends, and one already announced
stays known. The diff logic is reused from the health monitor rather than
copied. Maintenance page text updated.

Known limit: for Gitea this follows the repo's most recent run on any
workflow or branch, matching what the Gitea page shows; a failure in one
workflow can be masked by a later success of another.

Verified with 53 checks against fake Semaphore and Gitea servers and a
webhook receiver: classification, baseline (including not being consumed
when nothing is readable), single alert per failure, no repeat, in-progress/
stopped/cancelled runs, recovery and re-failure, unreadable project,
unreadable integration, run-fetch errors, maintenance windows (silenced,
then announced after), the toggle, disabled integrations and repos without
Actions. Real dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-26 03:11:59 +02:00
co-authored by Claude Sonnet 5
parent 9f1609c4ed
commit 4118062405
11 changed files with 219 additions and 5 deletions
+13
View File
@@ -148,6 +148,19 @@ which can be filtered by one or several tags (the filter is in the URL, so a
tag on a server's page links to everything sharing it), and they're searchable
from the global search box.
**Automation failures** — every 15 minutes the app looks at the latest run of
each Semaphore template and each Gitea repo's latest workflow run. A failed one
raises a single notification (with the project/template or repo, run number, and
for Gitea the run's link), and another when a later run succeeds. It's
state-based, so a job that fails every night alerts on the first failure rather
than every night. A run that's still going, was cancelled, or was stopped by
hand leaves things as they were, and anything that couldn't be read (a Semaphore
project or Gitea repo that errored, or an integration that's down) is neither
cleared nor re-announced. The first check after upgrading only records what's
already failing, so old failures aren't announced. Toggle it under Settings →
Notifications. For Gitea this follows the repo's most recent run on any
workflow or branch, the same as the Gitea page shows.
**Maintenance mode** silences alerts about one server, integration, or DNS
provider while you work on it (server offline / disk, storage and Synology
health, Proxmox backup alerts, and "integration down" for that service type).
+11 -3
View File
@@ -48,6 +48,8 @@ export interface SemaphoreTemplate {
export interface SemaphoreAdapter {
ping(): Promise<{ ok: boolean; latencyMs?: number; error?: string }>;
listTemplatesWithStatus(): Promise<SemaphoreTemplate[]>;
/** Like listTemplatesWithStatus, but says which projects couldn't be read — for callers that must not mistake "couldn't read" for "no templates". */
checkTemplates(): Promise<{ templates: SemaphoreTemplate[]; failedProjectIds: number[] }>;
runTemplate(projectId: number, templateId: number): Promise<SemaphoreTask>;
}
@@ -112,18 +114,24 @@ export function createSemaphoreAdapter(config: SemaphoreConfig): SemaphoreAdapte
}));
}
async function listTemplatesWithStatus(): Promise<SemaphoreTemplate[]> {
async function checkTemplates(): Promise<{ templates: SemaphoreTemplate[]; failedProjectIds: number[] }> {
const projects = await listProjects();
const failedProjectIds: number[] = [];
const perProject = await Promise.all(
projects.map(async (p) => {
try {
return await listTemplatesForProject(p.id, p.name);
} catch {
failedProjectIds.push(p.id);
return [];
}
}),
);
return perProject.flat();
return { templates: perProject.flat(), failedProjectIds };
}
async function listTemplatesWithStatus(): Promise<SemaphoreTemplate[]> {
return (await checkTemplates()).templates;
}
async function runTemplate(projectId: number, templateId: number): Promise<SemaphoreTask> {
@@ -141,5 +149,5 @@ export function createSemaphoreAdapter(config: SemaphoreConfig): SemaphoreAdapte
}
}
return withDiagLogging("semaphore", { ping, listTemplatesWithStatus, runTemplate });
return withDiagLogging("semaphore", { ping, listTemplatesWithStatus, checkTemplates, runTemplate });
}
+1
View File
@@ -69,6 +69,7 @@ const updateSchema = z.object({
dockerUpdateCheck: z.boolean(),
proxmoxBackupCheck: z.boolean(),
healthAlerts: z.boolean(),
automationAlerts: z.boolean(),
secretCheckTime: z.string().regex(/^\d{2}:\d{2}$/),
timezone: z.string(),
integrationFailureAlerts: z.boolean(),
+158
View File
@@ -0,0 +1,158 @@
import { and, eq, inArray } from "drizzle-orm";
import { db } from "../db/client.js";
import { integrations } from "../db/schema.js";
import { loadIntegrationConfig } from "../integrations/loadIntegration.js";
import { createSemaphoreAdapter } from "../integrations/semaphore/adapter.js";
import { createGiteaAdapter } from "../integrations/gitea/adapter.js";
import { diffConditions, type ActiveState, type HealthCondition } from "./healthMonitor.js";
import { activeSubjects } from "./maintenance.js";
import { notifyAutomationFailed, notifyAutomationRecovered } from "./notify.js";
import { getInternalFlag, setInternalFlag } from "./settingsStore.js";
const STATE_FLAG = "automationActiveConditions";
/**
* What a piece of automation's most recent run tells us. "unknown" covers everything that is neither a clear
* pass nor a clear failure — still running, waiting, cancelled, skipped, stopped by hand — and means "don't
* change what we were saying": a run in progress must not clear a failure it hasn't yet fixed, and a manual
* stop isn't a failure.
*/
export type RunOutcome = "failed" | "passed" | "unknown";
export interface AutomationItem {
/** Stable identity: the same template/repo always has the same key. */
key: string;
/** Hierarchical source, "semaphore:3:12:45" — held when the run's outcome isn't known. */
source: string;
outcome: RunOutcome;
/** What to say when it has failed. */
message: string;
}
// ─── Classification (pure) ──────────────────────────────────────────────────
export function classifySemaphoreStatus(status: string | null | undefined): RunOutcome {
if (status === "error") return "failed";
if (status === "success") return "passed";
return "unknown"; // waiting, starting, running, stopping, stopped, rejected, confirmed, ...
}
export function classifyGiteaRun(run: { status: string; conclusion: string | null }): RunOutcome {
if (run.conclusion === "failure" || run.status === "failure") return "failed";
if (run.conclusion === "success" || run.status === "success") return "passed";
return "unknown"; // running, waiting, blocked, cancelled, skipped, ...
}
function endedSuffix(end: string | null): string {
if (!end) return "";
const t = Date.parse(end);
return Number.isFinite(t) ? ` (${new Date(t).toISOString().slice(0, 16).replace("T", " ")} UTC)` : "";
}
/** Splits the items into the failures to report and the sources whose state must be left alone this pass. */
export function evaluateAutomation(items: AutomationItem[]): { conditions: HealthCondition[]; held: Set<string> } {
const conditions: HealthCondition[] = [];
const held = new Set<string>();
for (const item of items) {
if (item.outcome === "failed") conditions.push({ key: item.key, source: item.source, message: item.message });
else if (item.outcome === "unknown") held.add(item.source);
}
return { conditions, held };
}
// ─── Collection (I/O) ───────────────────────────────────────────────────────
async function collect(): Promise<{ items: AutomationItem[]; held: Set<string>; readable: number }> {
const items: AutomationItem[] = [];
const held = new Set<string>();
let readable = 0;
const rows = await db
.select({ id: integrations.id, name: integrations.name, type: integrations.type })
.from(integrations)
.where(and(eq(integrations.enabled, true), inArray(integrations.type, ["semaphore", "gitea"])));
for (const row of rows) {
try {
const loaded = await loadIntegrationConfig(row.id);
if (!loaded) continue;
if (row.type === "semaphore") {
const { templates, failedProjectIds } = await createSemaphoreAdapter(loaded.config as any).checkTemplates();
// A project that couldn't be listed contributes no templates — that's "couldn't read", not "all fixed".
for (const pid of failedProjectIds) held.add(`semaphore:${row.id}:${pid}`);
for (const t of templates) {
if (!t.lastTask) continue;
items.push({
key: `automation:semaphore:${row.id}:${t.projectId}:${t.id}`,
source: `semaphore:${row.id}:${t.projectId}:${t.id}`,
outcome: classifySemaphoreStatus(t.lastTask.status),
message: `${row.name} / ${t.projectName} / ${t.name}: run #${t.lastTask.id} failed${endedSuffix(t.lastTask.end)}`,
});
}
} else {
const repos = await createGiteaAdapter(loaded.config as any).listReposWithStatus();
for (const r of repos) {
if (!r.hasActions) continue;
const source = `gitea:${row.id}:${r.fullName}`;
// The adapter reports a run it couldn't fetch as null, the same as "no runs yet" — either way there's nothing to judge.
if (!r.latestRun) {
held.add(source);
continue;
}
const run = r.latestRun;
items.push({
key: `automation:gitea:${row.id}:${r.fullName}`,
source,
outcome: classifyGiteaRun(run),
message: `${row.name} / ${r.fullName}: "${run.displayTitle || "workflow"}" (run #${run.runNumber}) failed${run.headBranch ? ` on ${run.headBranch}` : ""}${run.htmlUrl ? ` — ${run.htmlUrl}` : ""}`,
});
}
}
readable++;
} catch (err) {
console.error(`[automation] couldn't read ${row.type} integration ${row.id}:`, err instanceof Error ? err.message : err);
held.add(`${row.type}:${row.id}`);
}
}
return { items, held, readable };
}
// ─── The scheduled pass ─────────────────────────────────────────────────────
async function loadState(): Promise<ActiveState | null> {
const raw = await getInternalFlag(STATE_FLAG);
if (raw === null) return null;
try {
return JSON.parse(raw);
} catch {
return {};
}
}
/**
* Reports each template/repo whose latest run failed, once, and again when it succeeds. State-based like the
* health monitor: a template that keeps failing every night alerts on the first failure, not every night.
*
* The very first pass only records what is already failing, without announcing it — on a fresh install (or
* the first time this feature runs) that would otherwise be a wall of alerts about failures that are months old.
*/
export async function runAutomationCheck(now: number = Date.now()): Promise<{ added: number; resolved: number; baseline: boolean }> {
const { items, held: readHeld, readable } = await collect();
const stored = await loadState();
// Nothing could be read at all (or nothing is configured): don't spend the "first pass" on an empty picture.
if (stored === null && readable === 0) return { added: 0, resolved: 0, baseline: false };
const { conditions, held: unknownHeld } = evaluateAutomation(items);
const held = new Set([...readHeld, ...unknownHeld]);
const { added, resolved, next } = diffConditions(stored ?? {}, conditions, held, await activeSubjects(new Date(now)));
await setInternalFlag(STATE_FLAG, JSON.stringify(next));
if (stored === null) return { added: 0, resolved: 0, baseline: true };
// State is tracked even with the alert toggle off (the notify functions check it), so turning it back on isn't a flood.
if (added.length > 0) await notifyAutomationFailed(added);
if (resolved.length > 0) await notifyAutomationRecovered(resolved);
return { added: added.length, resolved: resolved.length, baseline: false };
}
+9 -2
View File
@@ -1,4 +1,5 @@
import { runHealthCheck } from "./healthMonitor.js";
import { runAutomationCheck } from "./automationMonitor.js";
// 15 minutes matches the agent's default report interval, so a check never sees "stale" data that's just waiting for the next report.
const INTERVAL_MS = 15 * 60 * 1000;
@@ -6,17 +7,23 @@ const INTERVAL_MS = 15 * 60 * 1000;
let timer: ReturnType<typeof setInterval> | null = null;
async function pass(): Promise<void> {
// Independent, so a failure in one never skips the other.
try {
await runHealthCheck();
} catch (err) {
console.error("[health] check failed:", err);
}
try {
await runAutomationCheck();
} catch (err) {
console.error("[automation] check failed:", err);
}
}
export async function initHealthScheduler(): Promise<void> {
if (timer) return;
timer = setInterval(() => void pass(), INTERVAL_MS);
// Not awaited: the first pass talks to Proxmox/Synology and must not delay server startup.
// Not awaited: the first pass talks to Proxmox/Synology/Semaphore/Gitea and must not delay server startup.
void pass();
console.log(`[health] checking server/storage health every ${INTERVAL_MS / 60000} minutes`);
console.log(`[health] checking server/storage health and automation runs every ${INTERVAL_MS / 60000} minutes`);
}
+2
View File
@@ -22,6 +22,8 @@ export function subjectOfConditionKey(key: string): string | null {
if (m) return `integration:${m[1]}`;
m = /^(?:synology-volume|synology-disk|disk:synology):(\d+):/.exec(key);
if (m) return `integration:${m[1]}`;
m = /^automation:(?:semaphore|gitea):(\d+):/.exec(key);
if (m) return `integration:${m[1]}`;
return null;
}
+18
View File
@@ -295,6 +295,24 @@ export async function notifyHealthRecovered(resolved: { message: string }[]): Pr
);
}
export async function notifyAutomationFailed(issues: { message: string }[]): Promise<void> {
if (issues.length === 0) return;
if (!(await eventEnabled("automationAlerts"))) return;
await notify(
"Homelab Manager — Automation Failed",
`${issues.length} run${issues.length !== 1 ? "s" : ""} failed:\n\n${issues.map((i) => `✕ ${i.message}`).join("\n")}`,
);
}
export async function notifyAutomationRecovered(resolved: { message: string }[]): Promise<void> {
if (resolved.length === 0) return;
if (!(await eventEnabled("automationAlerts"))) return;
await notify(
"Homelab Manager — Automation Recovered",
`${resolved.length} no longer failing:\n\n${resolved.map((i) => `✓ ${i.message}`).join("\n")}`,
);
}
export async function notifyTailscaleKeyExpiry(
expiring: { integrationName: string; deviceLabel: string; daysLeft: number }[],
): Promise<void> {
+3
View File
@@ -43,6 +43,8 @@ export interface NotificationEvents {
dockerUpdateCheck: boolean;
proxmoxBackupCheck: boolean;
healthAlerts: boolean;
/** A Semaphore template or Gitea repo whose latest run failed, and when it succeeds again. */
automationAlerts: boolean;
secretCheckTime: string; // "HH:MM" — shared by the secret-expiry, Tailscale key-expiry, Docker update, and Proxmox backup checks
timezone: string;
integrationFailureAlerts: boolean;
@@ -114,6 +116,7 @@ const DEFAULTS: AppSettings = {
dockerUpdateCheck: true,
proxmoxBackupCheck: true,
healthAlerts: true,
automationAlerts: true,
secretCheckTime: "08:00",
timezone: "UTC",
integrationFailureAlerts: true,
+1
View File
@@ -183,6 +183,7 @@ export interface NotificationEvents {
dockerUpdateCheck: boolean;
proxmoxBackupCheck: boolean;
healthAlerts: boolean;
automationAlerts: boolean;
secretCheckTime: string;
timezone: string;
integrationFailureAlerts: boolean;
+1
View File
@@ -221,6 +221,7 @@ export default function Maintenance({ user }: { user: CurrentUser }) {
<li>Server offline and disk-full alerts (for a server).</li>
<li>Storage, volume and disk-health alerts (for a Proxmox or Synology integration).</li>
<li>Backup-failure and uncovered-guest alerts (for a Proxmox integration).</li>
<li>Failed-run alerts for the Semaphore or Gitea integration you choose.</li>
<li>
"Integration down" alerts for that <em>type</em> of service — these are tracked per type (e.g. all
Proxmox), not per instance, so with two Proxmox integrations a failure on the other one is silenced
@@ -47,6 +47,7 @@ const DEFAULT_NOTIFICATIONS: NotificationEvents = {
dockerUpdateCheck: true,
proxmoxBackupCheck: true,
healthAlerts: true,
automationAlerts: true,
secretCheckTime: "08:00",
timezone: "UTC",
integrationFailureAlerts: true,
@@ -526,6 +527,7 @@ export default function NotificationSettings() {
{ key: "dockerUpdateCheck" as const, label: "Docker image update available" },
{ key: "proxmoxBackupCheck" as const, label: "Proxmox backup failed or a guest has no coverage" },
{ key: "healthAlerts" as const, label: "Server offline, disk nearly full, or Synology volume/disk problem" },
{ key: "automationAlerts" as const, label: "Semaphore template or Gitea workflow run failed" },
{ key: "integrationFailureAlerts" as const, label: "Integration/DNS provider failing repeatedly" },
].map(({ key, label }) => (
<label key={key} className="form-check mb-2">