Files
Homelab-manager/server/src/services/maintenance.ts
T
bobbanandClaude Sonnet 5 4118062405 Alert when a Semaphore template or Gitea workflow run fails
Every 15 minutes (on the existing health-check timer) the app reads the
latest run of each Semaphore template and each Gitea repo's latest
workflow run. A failed one raises one notification, and another when a
later run succeeds. It is state-based like the server health alerts, so a
job that fails every night alerts on the first failure, not every night.
Gitea alerts include the run's link. Toggle: Settings > Notifications.

What counts:
- Semaphore "error" is a failure, "success" is a pass. A run that is
  waiting, running, stopped by hand or rejected is neither, so it leaves
  the previous state alone: a run in progress must not clear a failure it
  hasn't fixed yet, and a manual stop isn't a failure.
- Gitea failure/success likewise; running, waiting, blocked, cancelled and
  skipped leave things as they were.
- A failing template that gets another failing run does not re-alert.

Not mistaking "couldn't read" for "fixed":
- Semaphore's template listing swallowed per-project errors, so a project
  that failed to load looked like a project with no templates. A new
  checkTemplates adapter method reports which projects failed, and their
  failures are held rather than cleared.
- Gitea reports a run it couldn't fetch as null, the same as "no runs";
  both leave the repo's state alone.
- An unreachable integration holds all of its failures. Nothing is cleared
  or re-announced while it is down.

The first pass only records what is already failing without announcing
it, so upgrading (or adding an integration to a fresh install) doesn't
produce a wall of alerts about months-old failures. That baseline is not
spent while nothing could be read.

Maintenance windows on a Semaphore or Gitea integration silence its
failure alerts with the same rules as the health alerts: a problem that
starts during a window alerts when it ends, and one already announced
stays known. The diff logic is reused from the health monitor rather than
copied. Maintenance page text updated.

Known limit: for Gitea this follows the repo's most recent run on any
workflow or branch, matching what the Gitea page shows; a failure in one
workflow can be masked by a later success of another.

Verified with 53 checks against fake Semaphore and Gitea servers and a
webhook receiver: classification, baseline (including not being consumed
when nothing is readable), single alert per failure, no repeat, in-progress/
stopped/cancelled runs, recovery and re-failure, unreadable project,
unreadable integration, run-fetch errors, maintenance windows (silenced,
then announced after), the toggle, disabled integrations and repos without
Actions. Real dev database mtime untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 03:11:59 +02:00

74 lines
3.7 KiB
TypeScript

import { eq, gt } from "drizzle-orm";
import { db } from "../db/client.js";
import { dnsProviders, integrations, maintenanceWindows, servers } from "../db/schema.js";
export type ActiveWindow = typeof maintenanceWindows.$inferSelect;
/** Windows whose end is still in the future. ISO strings compare correctly as text, so this is a plain string comparison. */
export async function listActiveWindows(now: Date = new Date()): Promise<ActiveWindow[]> {
return db.select().from(maintenanceWindows).where(gt(maintenanceWindows.endsAt, now.toISOString()));
}
/**
* The identity a health condition or alert belongs to, as "server:3" /
* "integration:1" / "dns_provider:2", so it can be matched against active windows.
* Derived from the condition key (which already encodes it) rather than stored,
* so conditions persisted by an earlier version still map correctly.
*/
export function subjectOfConditionKey(key: string): string | null {
let m = /^(?:offline|disk):server:(\d+)(?::|$)/.exec(key);
if (m) return `server:${m[1]}`;
m = /^disk:proxmox:(\d+):/.exec(key);
if (m) return `integration:${m[1]}`;
m = /^(?:synology-volume|synology-disk|disk:synology):(\d+):/.exec(key);
if (m) return `integration:${m[1]}`;
m = /^automation:(?:semaphore|gitea):(\d+):/.exec(key);
if (m) return `integration:${m[1]}`;
return null;
}
export async function activeSubjects(now: Date = new Date()): Promise<Set<string>> {
return new Set((await listActiveWindows(now)).map((w) => `${w.targetType}:${w.targetId}`));
}
export async function isInMaintenance(targetType: ActiveWindow["targetType"], targetId: number, now: Date = new Date()): Promise<boolean> {
return (await activeSubjects(now)).has(`${targetType}:${targetId}`);
}
/**
* Integration-failure alerts are tracked per service TYPE ("proxmox",
* "cloudflare", ...), not per configured instance, so a window on any
* integration or DNS provider of that type silences that type's failure alerts.
* (With two integrations of one type, a real failure on the un-windowed one is
* silenced too while the window is open — a known consequence of that
* granularity, stated on the Maintenance page.)
*/
export async function isSourceInMaintenance(source: string, now: Date = new Date()): Promise<boolean> {
const windows = await listActiveWindows(now);
if (windows.length === 0) return false;
for (const w of windows) {
if (w.targetType === "integration") {
const [row] = await db.select({ type: integrations.type }).from(integrations).where(eq(integrations.id, w.targetId)).limit(1);
if (row?.type === source) return true;
} else if (w.targetType === "dns_provider") {
const [row] = await db.select({ type: dnsProviders.providerType }).from(dnsProviders).where(eq(dnsProviders.id, w.targetId)).limit(1);
if (row?.type === source) return true;
}
}
return false;
}
/** Display name for a window's target, or null if the target no longer exists. */
export async function describeTarget(targetType: ActiveWindow["targetType"], targetId: number): Promise<{ name: string; kind: string } | null> {
if (targetType === "server") {
const [r] = await db.select({ name: servers.name }).from(servers).where(eq(servers.id, targetId)).limit(1);
return r ? { name: r.name, kind: "Server" } : null;
}
if (targetType === "integration") {
const [r] = await db.select({ name: integrations.name, type: integrations.type }).from(integrations).where(eq(integrations.id, targetId)).limit(1);
return r ? { name: r.name, kind: `Integration (${r.type})` } : null;
}
const [r] = await db.select({ name: dnsProviders.name, type: dnsProviders.providerType }).from(dnsProviders).where(eq(dnsProviders.id, targetId)).limit(1);
return r ? { name: r.name, kind: `DNS provider (${r.type})` } : null;
}