Add an Alerts page under Operations listing everything that's wrong now

One list of the current problems across servers and integrations, instead
of waiting for a notification or visiting each page: servers that stopped
reporting, full or nearly full disks and volumes (critical from 95%),
Synology volume/disk problems, failed or uncovered Proxmox backups,
failed Proxmox Backup Server verifications, container image updates,
expired or expiring secrets/domains/Tailscale keys, failed Semaphore and
Gitea runs, Uptime Kuma monitors that are down, overdue osTicket tickets,
and integrations whose calls keep failing. Visible to every role, with
severity and kind filters, search, sorting, CSV export and "Check now".

It runs the same checks that send the notifications rather than a second
copy of them: the detection in the health, automation, Proxmox backup,
PBS, Docker update and Tailscale key checks is pulled out into shared
collectors that both the schedulers and the page call, so the two can't
disagree about what counts as a problem. Notification behaviour is
unchanged, including the scheduled backup checks skipping integrations
under a maintenance window. Unlike the notifications the page ignores the
on/off toggles, and keeps problems under a maintenance window, marked
silenced and counted apart.

It reads live, so a result is reused for a minute (and Refresh can't
re-run everything more than once every ten seconds), and every source has
a 20 s limit so one hung integration can't hang the page. Anything it
couldn't read is called out at the top instead of looking like all clear,
and server checks pause for the same 20 minutes after a restart as the
notifications do, with a note saying so.

Also gives the newer integrations (PBS, osTicket, Uptime Kuma, phpIPAM)
proper names in "integration down" notifications instead of their ids.

Verified through the real routes against a scratch database with fake
backends (offline and full-disk servers, secrets and domains, a silenced
server, a fake PBS with failed verification, a hanging integration, a
refused one, a failing-calls streak, caching, the restart grace period,
auth), and by rendering the real page against that data in a browser:
filters, search, silenced toggle, sorting, Check now, dark mode.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
bobbanandClaude Sonnet 5.5 committed 2026-10-02 23:07:57 +02:00
1 parent ad1fb5338f
commit 447f33fff6
19 files changed
+883 -28

No files matched your search

+15
View File
@@ -255,6 +255,21 @@ already failing, so old failures aren't announced. Toggle it under Settings →
Notifications. For Gitea this follows the repo's most recent run on any
workflow or branch, the same as the Gitea page shows.
**Alerts** (Operations → Alerts, visible to every role) lists everything that's
wrong right now in one place, instead of waiting for a notification or visiting
each page: servers that stopped reporting, full or nearly full disks and volumes
(critical from 95%), Synology volume/disk problems, failed or uncovered Proxmox
backups, failed Proxmox Backup Server verifications, container image updates,
secrets, domains and Tailscale keys that are expired or about to be, failed
Semaphore/Gitea runs, Uptime Kuma monitors that are down, overdue osTicket
tickets, and integrations whose calls keep failing. It runs the same checks that
send the notifications — so the two can't disagree — but ignores the on/off
toggles, since it's for looking at rather than being interrupted by. Problems
under a maintenance window stay listed, marked silenced and counted separately.
It checks live (a recent result is reused for a minute; "Check now" forces a
fresh one), and anything it couldn't read is called out at the top rather than
quietly treated as fine.
**Maintenance mode** silences alerts about one server, integration, or DNS
provider while you work on it (server offline / disk, storage and Synology
health, Proxmox backup alerts, Proxmox Backup Server verification alerts, and