Add an Alerts page under Operations listing everything that's wrong now

One list of the current problems across servers and integrations, instead
of waiting for a notification or visiting each page: servers that stopped
reporting, full or nearly full disks and volumes (critical from 95%),
Synology volume/disk problems, failed or uncovered Proxmox backups,
failed Proxmox Backup Server verifications, container image updates,
expired or expiring secrets/domains/Tailscale keys, failed Semaphore and
Gitea runs, Uptime Kuma monitors that are down, overdue osTicket tickets,
and integrations whose calls keep failing. Visible to every role, with
severity and kind filters, search, sorting, CSV export and "Check now".

It runs the same checks that send the notifications rather than a second
copy of them: the detection in the health, automation, Proxmox backup,
PBS, Docker update and Tailscale key checks is pulled out into shared
collectors that both the schedulers and the page call, so the two can't
disagree about what counts as a problem. Notification behaviour is
unchanged, including the scheduled backup checks skipping integrations
under a maintenance window. Unlike the notifications the page ignores the
on/off toggles, and keeps problems under a maintenance window, marked
silenced and counted apart.

It reads live, so a result is reused for a minute (and Refresh can't
re-run everything more than once every ten seconds), and every source has
a 20 s limit so one hung integration can't hang the page. Anything it
couldn't read is called out at the top instead of looking like all clear,
and server checks pause for the same 20 minutes after a restart as the
notifications do, with a note saying so.

Also gives the newer integrations (PBS, osTicket, Uptime Kuma, phpIPAM)
proper names in "integration down" notifications instead of their ids.

Verified through the real routes against a scratch database with fake
backends (offline and full-disk servers, secrets and domains, a silenced
server, a fake PBS with failed verification, a hanging integration, a
refused one, a failing-calls streak, caching, the restart grace period,
auth), and by rendering the real page against that data in a browser:
filters, search, silenced toggle, sorting, Check now, dark mode.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
bobbanandClaude Sonnet 5.5 committed 2026-10-02 23:07:57 +02:00
1 parent ad1fb5338f
commit 447f33fff6
19 files changed
+883 -28

No files matched your search

+14
View File
@@ -89,6 +89,20 @@ schedule.
|---|---|
| "Notifications from quiet hours" | Once, at quiet hours' configured end time, **only if enabled and only if at least one notification was held** during the window. Bundles every held notification's title and message into one message, then clears the queue. |
## The Alerts page
**Operations → Alerts** shows what these notifications are about — the problems that exist *right now* — as a list you can look at, filter, and
export. It uses the same checks as the notifications above, so a problem appears there for exactly the reason it would be notified, but it differs
in three ways:
- It ignores the "Notify on" toggles. Turning a notification off doesn't hide the problem from the page.
- Problems under a **maintenance window** are kept on the list, marked *silenced* and counted separately, instead of being dropped.
- It also lists things nothing notifies about: Uptime Kuma monitors that are down, osTicket tickets that are overdue, and any integration that's
failing its last few calls (before the threshold that triggers an "integration down" notification).
It runs the checks live when opened (a recent result is reused for a minute), and shows what it couldn't read at the top, so a missing section means
"couldn't check" and not "all clear".
## What does *not* send a notification
Worth calling out explicitly, since it's easy to assume everything in