Alert when a Semaphore template or Gitea workflow run fails
Every 15 minutes (on the existing health-check timer) the app reads the latest run of each Semaphore template and each Gitea repo's latest workflow run. A failed one raises one notification, and another when a later run succeeds. It is state-based like the server health alerts, so a job that fails every night alerts on the first failure, not every night. Gitea alerts include the run's link. Toggle: Settings > Notifications. What counts: - Semaphore "error" is a failure, "success" is a pass. A run that is waiting, running, stopped by hand or rejected is neither, so it leaves the previous state alone: a run in progress must not clear a failure it hasn't fixed yet, and a manual stop isn't a failure. - Gitea failure/success likewise; running, waiting, blocked, cancelled and skipped leave things as they were. - A failing template that gets another failing run does not re-alert. Not mistaking "couldn't read" for "fixed": - Semaphore's template listing swallowed per-project errors, so a project that failed to load looked like a project with no templates. A new checkTemplates adapter method reports which projects failed, and their failures are held rather than cleared. - Gitea reports a run it couldn't fetch as null, the same as "no runs"; both leave the repo's state alone. - An unreachable integration holds all of its failures. Nothing is cleared or re-announced while it is down. The first pass only records what is already failing without announcing it, so upgrading (or adding an integration to a fresh install) doesn't produce a wall of alerts about months-old failures. That baseline is not spent while nothing could be read. Maintenance windows on a Semaphore or Gitea integration silence its failure alerts with the same rules as the health alerts: a problem that starts during a window alerts when it ends, and one already announced stays known. The diff logic is reused from the health monitor rather than copied. Maintenance page text updated. Known limit: for Gitea this follows the repo's most recent run on any workflow or branch, matching what the Gitea page shows; a failure in one workflow can be masked by a later success of another. Verified with 53 checks against fake Semaphore and Gitea servers and a webhook receiver: classification, baseline (including not being consumed when nothing is readable), single alert per failure, no repeat, in-progress/ stopped/cancelled runs, recovery and re-failure, unreadable project, unreadable integration, run-fetch errors, maintenance windows (silenced, then announced after), the toggle, disabled integrations and repos without Actions. Real dev database mtime untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
1 parent
9f1609c4ed
commit
4118062405
11 files changed
+219
-5
No files matched your search
@@ -148,6 +148,19 @@ which can be filtered by one or several tags (the filter is in the URL, so a
|
||||
tag on a server's page links to everything sharing it), and they're searchable
|
||||
from the global search box.
|
||||
|
||||
**Automation failures** — every 15 minutes the app looks at the latest run of
|
||||
each Semaphore template and each Gitea repo's latest workflow run. A failed one
|
||||
raises a single notification (with the project/template or repo, run number, and
|
||||
for Gitea the run's link), and another when a later run succeeds. It's
|
||||
state-based, so a job that fails every night alerts on the first failure rather
|
||||
than every night. A run that's still going, was cancelled, or was stopped by
|
||||
hand leaves things as they were, and anything that couldn't be read (a Semaphore
|
||||
project or Gitea repo that errored, or an integration that's down) is neither
|
||||
cleared nor re-announced. The first check after upgrading only records what's
|
||||
already failing, so old failures aren't announced. Toggle it under Settings →
|
||||
Notifications. For Gitea this follows the repo's most recent run on any
|
||||
workflow or branch, the same as the Gitea page shows.
|
||||
|
||||
**Maintenance mode** silences alerts about one server, integration, or DNS
|
||||
provider while you work on it (server offline / disk, storage and Synology
|
||||
health, Proxmox backup alerts, and "integration down" for that service type).
|
||||
|
||||
@@ -48,6 +48,8 @@ export interface SemaphoreTemplate {
|
||||
export interface SemaphoreAdapter {
|
||||
ping(): Promise<{ ok: boolean; latencyMs?: number; error?: string }>;
|
||||
listTemplatesWithStatus(): Promise<SemaphoreTemplate[]>;
|
||||
/** Like listTemplatesWithStatus, but says which projects couldn't be read — for callers that must not mistake "couldn't read" for "no templates". */
|
||||
checkTemplates(): Promise<{ templates: SemaphoreTemplate[]; failedProjectIds: number[] }>;
|
||||
runTemplate(projectId: number, templateId: number): Promise<SemaphoreTask>;
|
||||
}
|
||||
|
||||
@@ -112,18 +114,24 @@ export function createSemaphoreAdapter(config: SemaphoreConfig): SemaphoreAdapte
|
||||
}));
|
||||
}
|
||||
|
||||
async function listTemplatesWithStatus(): Promise<SemaphoreTemplate[]> {
|
||||
async function checkTemplates(): Promise<{ templates: SemaphoreTemplate[]; failedProjectIds: number[] }> {
|
||||
const projects = await listProjects();
|
||||
const failedProjectIds: number[] = [];
|
||||
const perProject = await Promise.all(
|
||||
projects.map(async (p) => {
|
||||
try {
|
||||
return await listTemplatesForProject(p.id, p.name);
|
||||
} catch {
|
||||
failedProjectIds.push(p.id);
|
||||
return [];
|
||||
}
|
||||
}),
|
||||
);
|
||||
return perProject.flat();
|
||||
return { templates: perProject.flat(), failedProjectIds };
|
||||
}
|
||||
|
||||
async function listTemplatesWithStatus(): Promise<SemaphoreTemplate[]> {
|
||||
return (await checkTemplates()).templates;
|
||||
}
|
||||
|
||||
async function runTemplate(projectId: number, templateId: number): Promise<SemaphoreTask> {
|
||||
@@ -141,5 +149,5 @@ export function createSemaphoreAdapter(config: SemaphoreConfig): SemaphoreAdapte
|
||||
}
|
||||
}
|
||||
|
||||
return withDiagLogging("semaphore", { ping, listTemplatesWithStatus, runTemplate });
|
||||
return withDiagLogging("semaphore", { ping, listTemplatesWithStatus, checkTemplates, runTemplate });
|
||||
}
|
||||
@@ -69,6 +69,7 @@ const updateSchema = z.object({
|
||||
dockerUpdateCheck: z.boolean(),
|
||||
proxmoxBackupCheck: z.boolean(),
|
||||
healthAlerts: z.boolean(),
|
||||
automationAlerts: z.boolean(),
|
||||
secretCheckTime: z.string().regex(/^\d{2}:\d{2}$/),
|
||||
timezone: z.string(),
|
||||
integrationFailureAlerts: z.boolean(),
|
||||
|
||||
@@ -0,0 +1,158 @@
|
||||
import { and, eq, inArray } from "drizzle-orm";
|
||||
import { db } from "../db/client.js";
|
||||
import { integrations } from "../db/schema.js";
|
||||
import { loadIntegrationConfig } from "../integrations/loadIntegration.js";
|
||||
import { createSemaphoreAdapter } from "../integrations/semaphore/adapter.js";
|
||||
import { createGiteaAdapter } from "../integrations/gitea/adapter.js";
|
||||
import { diffConditions, type ActiveState, type HealthCondition } from "./healthMonitor.js";
|
||||
import { activeSubjects } from "./maintenance.js";
|
||||
import { notifyAutomationFailed, notifyAutomationRecovered } from "./notify.js";
|
||||
import { getInternalFlag, setInternalFlag } from "./settingsStore.js";
|
||||
|
||||
const STATE_FLAG = "automationActiveConditions";
|
||||
|
||||
/**
|
||||
* What a piece of automation's most recent run tells us. "unknown" covers everything that is neither a clear
|
||||
* pass nor a clear failure — still running, waiting, cancelled, skipped, stopped by hand — and means "don't
|
||||
* change what we were saying": a run in progress must not clear a failure it hasn't yet fixed, and a manual
|
||||
* stop isn't a failure.
|
||||
*/
|
||||
export type RunOutcome = "failed" | "passed" | "unknown";
|
||||
|
||||
export interface AutomationItem {
|
||||
/** Stable identity: the same template/repo always has the same key. */
|
||||
key: string;
|
||||
/** Hierarchical source, "semaphore:3:12:45" — held when the run's outcome isn't known. */
|
||||
source: string;
|
||||
outcome: RunOutcome;
|
||||
/** What to say when it has failed. */
|
||||
message: string;
|
||||
}
|
||||
|
||||
// ─── Classification (pure) ──────────────────────────────────────────────────
|
||||
|
||||
export function classifySemaphoreStatus(status: string | null | undefined): RunOutcome {
|
||||
if (status === "error") return "failed";
|
||||
if (status === "success") return "passed";
|
||||
return "unknown"; // waiting, starting, running, stopping, stopped, rejected, confirmed, ...
|
||||
}
|
||||
|
||||
export function classifyGiteaRun(run: { status: string; conclusion: string | null }): RunOutcome {
|
||||
if (run.conclusion === "failure" || run.status === "failure") return "failed";
|
||||
if (run.conclusion === "success" || run.status === "success") return "passed";
|
||||
return "unknown"; // running, waiting, blocked, cancelled, skipped, ...
|
||||
}
|
||||
|
||||
function endedSuffix(end: string | null): string {
|
||||
if (!end) return "";
|
||||
const t = Date.parse(end);
|
||||
return Number.isFinite(t) ? ` (${new Date(t).toISOString().slice(0, 16).replace("T", " ")} UTC)` : "";
|
||||
}
|
||||
|
||||
/** Splits the items into the failures to report and the sources whose state must be left alone this pass. */
|
||||
export function evaluateAutomation(items: AutomationItem[]): { conditions: HealthCondition[]; held: Set<string> } {
|
||||
const conditions: HealthCondition[] = [];
|
||||
const held = new Set<string>();
|
||||
for (const item of items) {
|
||||
if (item.outcome === "failed") conditions.push({ key: item.key, source: item.source, message: item.message });
|
||||
else if (item.outcome === "unknown") held.add(item.source);
|
||||
}
|
||||
return { conditions, held };
|
||||
}
|
||||
|
||||
// ─── Collection (I/O) ───────────────────────────────────────────────────────
|
||||
|
||||
async function collect(): Promise<{ items: AutomationItem[]; held: Set<string>; readable: number }> {
|
||||
const items: AutomationItem[] = [];
|
||||
const held = new Set<string>();
|
||||
let readable = 0;
|
||||
|
||||
const rows = await db
|
||||
.select({ id: integrations.id, name: integrations.name, type: integrations.type })
|
||||
.from(integrations)
|
||||
.where(and(eq(integrations.enabled, true), inArray(integrations.type, ["semaphore", "gitea"])));
|
||||
|
||||
for (const row of rows) {
|
||||
try {
|
||||
const loaded = await loadIntegrationConfig(row.id);
|
||||
if (!loaded) continue;
|
||||
|
||||
if (row.type === "semaphore") {
|
||||
const { templates, failedProjectIds } = await createSemaphoreAdapter(loaded.config as any).checkTemplates();
|
||||
// A project that couldn't be listed contributes no templates — that's "couldn't read", not "all fixed".
|
||||
for (const pid of failedProjectIds) held.add(`semaphore:${row.id}:${pid}`);
|
||||
for (const t of templates) {
|
||||
if (!t.lastTask) continue;
|
||||
items.push({
|
||||
key: `automation:semaphore:${row.id}:${t.projectId}:${t.id}`,
|
||||
source: `semaphore:${row.id}:${t.projectId}:${t.id}`,
|
||||
outcome: classifySemaphoreStatus(t.lastTask.status),
|
||||
message: `${row.name} / ${t.projectName} / ${t.name}: run #${t.lastTask.id} failed${endedSuffix(t.lastTask.end)}`,
|
||||
});
|
||||
}
|
||||
} else {
|
||||
const repos = await createGiteaAdapter(loaded.config as any).listReposWithStatus();
|
||||
for (const r of repos) {
|
||||
if (!r.hasActions) continue;
|
||||
const source = `gitea:${row.id}:${r.fullName}`;
|
||||
// The adapter reports a run it couldn't fetch as null, the same as "no runs yet" — either way there's nothing to judge.
|
||||
if (!r.latestRun) {
|
||||
held.add(source);
|
||||
continue;
|
||||
}
|
||||
const run = r.latestRun;
|
||||
items.push({
|
||||
key: `automation:gitea:${row.id}:${r.fullName}`,
|
||||
source,
|
||||
outcome: classifyGiteaRun(run),
|
||||
message: `${row.name} / ${r.fullName}: "${run.displayTitle || "workflow"}" (run #${run.runNumber}) failed${run.headBranch ? ` on ${run.headBranch}` : ""}${run.htmlUrl ? ` — ${run.htmlUrl}` : ""}`,
|
||||
});
|
||||
}
|
||||
}
|
||||
readable++;
|
||||
} catch (err) {
|
||||
console.error(`[automation] couldn't read ${row.type} integration ${row.id}:`, err instanceof Error ? err.message : err);
|
||||
held.add(`${row.type}:${row.id}`);
|
||||
}
|
||||
}
|
||||
return { items, held, readable };
|
||||
}
|
||||
|
||||
// ─── The scheduled pass ─────────────────────────────────────────────────────
|
||||
|
||||
async function loadState(): Promise<ActiveState | null> {
|
||||
const raw = await getInternalFlag(STATE_FLAG);
|
||||
if (raw === null) return null;
|
||||
try {
|
||||
return JSON.parse(raw);
|
||||
} catch {
|
||||
return {};
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Reports each template/repo whose latest run failed, once, and again when it succeeds. State-based like the
|
||||
* health monitor: a template that keeps failing every night alerts on the first failure, not every night.
|
||||
*
|
||||
* The very first pass only records what is already failing, without announcing it — on a fresh install (or
|
||||
* the first time this feature runs) that would otherwise be a wall of alerts about failures that are months old.
|
||||
*/
|
||||
export async function runAutomationCheck(now: number = Date.now()): Promise<{ added: number; resolved: number; baseline: boolean }> {
|
||||
const { items, held: readHeld, readable } = await collect();
|
||||
const stored = await loadState();
|
||||
|
||||
// Nothing could be read at all (or nothing is configured): don't spend the "first pass" on an empty picture.
|
||||
if (stored === null && readable === 0) return { added: 0, resolved: 0, baseline: false };
|
||||
|
||||
const { conditions, held: unknownHeld } = evaluateAutomation(items);
|
||||
const held = new Set([...readHeld, ...unknownHeld]);
|
||||
const { added, resolved, next } = diffConditions(stored ?? {}, conditions, held, await activeSubjects(new Date(now)));
|
||||
await setInternalFlag(STATE_FLAG, JSON.stringify(next));
|
||||
|
||||
if (stored === null) return { added: 0, resolved: 0, baseline: true };
|
||||
|
||||
// State is tracked even with the alert toggle off (the notify functions check it), so turning it back on isn't a flood.
|
||||
if (added.length > 0) await notifyAutomationFailed(added);
|
||||
if (resolved.length > 0) await notifyAutomationRecovered(resolved);
|
||||
return { added: added.length, resolved: resolved.length, baseline: false };
|
||||
}
|
||||
@@ -1,4 +1,5 @@
|
||||
import { runHealthCheck } from "./healthMonitor.js";
|
||||
import { runAutomationCheck } from "./automationMonitor.js";
|
||||
|
||||
// 15 minutes matches the agent's default report interval, so a check never sees "stale" data that's just waiting for the next report.
|
||||
const INTERVAL_MS = 15 * 60 * 1000;
|
||||
@@ -6,17 +7,23 @@ const INTERVAL_MS = 15 * 60 * 1000;
|
||||
let timer: ReturnType<typeof setInterval> | null = null;
|
||||
|
||||
async function pass(): Promise<void> {
|
||||
// Independent, so a failure in one never skips the other.
|
||||
try {
|
||||
await runHealthCheck();
|
||||
} catch (err) {
|
||||
console.error("[health] check failed:", err);
|
||||
}
|
||||
try {
|
||||
await runAutomationCheck();
|
||||
} catch (err) {
|
||||
console.error("[automation] check failed:", err);
|
||||
}
|
||||
}
|
||||
|
||||
export async function initHealthScheduler(): Promise<void> {
|
||||
if (timer) return;
|
||||
timer = setInterval(() => void pass(), INTERVAL_MS);
|
||||
// Not awaited: the first pass talks to Proxmox/Synology and must not delay server startup.
|
||||
// Not awaited: the first pass talks to Proxmox/Synology/Semaphore/Gitea and must not delay server startup.
|
||||
void pass();
|
||||
console.log(`[health] checking server/storage health every ${INTERVAL_MS / 60000} minutes`);
|
||||
console.log(`[health] checking server/storage health and automation runs every ${INTERVAL_MS / 60000} minutes`);
|
||||
}
|
||||
@@ -22,6 +22,8 @@ export function subjectOfConditionKey(key: string): string | null {
|
||||
if (m) return `integration:${m[1]}`;
|
||||
m = /^(?:synology-volume|synology-disk|disk:synology):(\d+):/.exec(key);
|
||||
if (m) return `integration:${m[1]}`;
|
||||
m = /^automation:(?:semaphore|gitea):(\d+):/.exec(key);
|
||||
if (m) return `integration:${m[1]}`;
|
||||
return null;
|
||||
}
|
||||
|
||||
|
||||
@@ -295,6 +295,24 @@ export async function notifyHealthRecovered(resolved: { message: string }[]): Pr
|
||||
);
|
||||
}
|
||||
|
||||
export async function notifyAutomationFailed(issues: { message: string }[]): Promise<void> {
|
||||
if (issues.length === 0) return;
|
||||
if (!(await eventEnabled("automationAlerts"))) return;
|
||||
await notify(
|
||||
"Homelab Manager — Automation Failed",
|
||||
`${issues.length} run${issues.length !== 1 ? "s" : ""} failed:\n\n${issues.map((i) => `✕ ${i.message}`).join("\n")}`,
|
||||
);
|
||||
}
|
||||
|
||||
export async function notifyAutomationRecovered(resolved: { message: string }[]): Promise<void> {
|
||||
if (resolved.length === 0) return;
|
||||
if (!(await eventEnabled("automationAlerts"))) return;
|
||||
await notify(
|
||||
"Homelab Manager — Automation Recovered",
|
||||
`${resolved.length} no longer failing:\n\n${resolved.map((i) => `✓ ${i.message}`).join("\n")}`,
|
||||
);
|
||||
}
|
||||
|
||||
export async function notifyTailscaleKeyExpiry(
|
||||
expiring: { integrationName: string; deviceLabel: string; daysLeft: number }[],
|
||||
): Promise<void> {
|
||||
|
||||
@@ -43,6 +43,8 @@ export interface NotificationEvents {
|
||||
dockerUpdateCheck: boolean;
|
||||
proxmoxBackupCheck: boolean;
|
||||
healthAlerts: boolean;
|
||||
/** A Semaphore template or Gitea repo whose latest run failed, and when it succeeds again. */
|
||||
automationAlerts: boolean;
|
||||
secretCheckTime: string; // "HH:MM" — shared by the secret-expiry, Tailscale key-expiry, Docker update, and Proxmox backup checks
|
||||
timezone: string;
|
||||
integrationFailureAlerts: boolean;
|
||||
@@ -114,6 +116,7 @@ const DEFAULTS: AppSettings = {
|
||||
dockerUpdateCheck: true,
|
||||
proxmoxBackupCheck: true,
|
||||
healthAlerts: true,
|
||||
automationAlerts: true,
|
||||
secretCheckTime: "08:00",
|
||||
timezone: "UTC",
|
||||
integrationFailureAlerts: true,
|
||||
|
||||
@@ -183,6 +183,7 @@ export interface NotificationEvents {
|
||||
dockerUpdateCheck: boolean;
|
||||
proxmoxBackupCheck: boolean;
|
||||
healthAlerts: boolean;
|
||||
automationAlerts: boolean;
|
||||
secretCheckTime: string;
|
||||
timezone: string;
|
||||
integrationFailureAlerts: boolean;
|
||||
|
||||
@@ -221,6 +221,7 @@ export default function Maintenance({ user }: { user: CurrentUser }) {
|
||||
<li>Server offline and disk-full alerts (for a server).</li>
|
||||
<li>Storage, volume and disk-health alerts (for a Proxmox or Synology integration).</li>
|
||||
<li>Backup-failure and uncovered-guest alerts (for a Proxmox integration).</li>
|
||||
<li>Failed-run alerts for the Semaphore or Gitea integration you choose.</li>
|
||||
<li>
|
||||
"Integration down" alerts for that <em>type</em> of service — these are tracked per type (e.g. all
|
||||
Proxmox), not per instance, so with two Proxmox integrations a failure on the other one is silenced
|
||||
|
||||
@@ -47,6 +47,7 @@ const DEFAULT_NOTIFICATIONS: NotificationEvents = {
|
||||
dockerUpdateCheck: true,
|
||||
proxmoxBackupCheck: true,
|
||||
healthAlerts: true,
|
||||
automationAlerts: true,
|
||||
secretCheckTime: "08:00",
|
||||
timezone: "UTC",
|
||||
integrationFailureAlerts: true,
|
||||
@@ -526,6 +527,7 @@ export default function NotificationSettings() {
|
||||
{ key: "dockerUpdateCheck" as const, label: "Docker image update available" },
|
||||
{ key: "proxmoxBackupCheck" as const, label: "Proxmox backup failed or a guest has no coverage" },
|
||||
{ key: "healthAlerts" as const, label: "Server offline, disk nearly full, or Synology volume/disk problem" },
|
||||
{ key: "automationAlerts" as const, label: "Semaphore template or Gitea workflow run failed" },
|
||||
{ key: "integrationFailureAlerts" as const, label: "Integration/DNS provider failing repeatedly" },
|
||||
].map(({ key, label }) => (
|
||||
<label key={key} className="form-check mb-2">
|
||||
|
||||
Reference in new issue
Block a user