Add an Alerts page under Operations listing everything that's wrong now
One list of the current problems across servers and integrations, instead of waiting for a notification or visiting each page: servers that stopped reporting, full or nearly full disks and volumes (critical from 95%), Synology volume/disk problems, failed or uncovered Proxmox backups, failed Proxmox Backup Server verifications, container image updates, expired or expiring secrets/domains/Tailscale keys, failed Semaphore and Gitea runs, Uptime Kuma monitors that are down, overdue osTicket tickets, and integrations whose calls keep failing. Visible to every role, with severity and kind filters, search, sorting, CSV export and "Check now". It runs the same checks that send the notifications rather than a second copy of them: the detection in the health, automation, Proxmox backup, PBS, Docker update and Tailscale key checks is pulled out into shared collectors that both the schedulers and the page call, so the two can't disagree about what counts as a problem. Notification behaviour is unchanged, including the scheduled backup checks skipping integrations under a maintenance window. Unlike the notifications the page ignores the on/off toggles, and keeps problems under a maintenance window, marked silenced and counted apart. It reads live, so a result is reused for a minute (and Refresh can't re-run everything more than once every ten seconds), and every source has a 20 s limit so one hung integration can't hang the page. Anything it couldn't read is called out at the top instead of looking like all clear, and server checks pause for the same 20 minutes after a restart as the notifications do, with a note saying so. Also gives the newer integrations (PBS, osTicket, Uptime Kuma, phpIPAM) proper names in "integration down" notifications instead of their ids. Verified through the real routes against a scratch database with fake backends (offline and full-disk servers, secrets and domains, a silenced server, a fake PBS with failed verification, a hanging integration, a refused one, a failing-calls streak, caching, the restart grace period, auth), and by rendering the real page against that data in a browser: filters, search, silenced toggle, sorting, Check now, dark mode. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
1 parent
ad1fb5338f
commit
447f33fff6
19 files changed
+883
-28
No files matched your search
@@ -89,6 +89,20 @@ schedule.
|
||||
|---|---|
|
||||
| "Notifications from quiet hours" | Once, at quiet hours' configured end time, **only if enabled and only if at least one notification was held** during the window. Bundles every held notification's title and message into one message, then clears the queue. |
|
||||
|
||||
## The Alerts page
|
||||
|
||||
**Operations → Alerts** shows what these notifications are about — the problems that exist *right now* — as a list you can look at, filter, and
|
||||
export. It uses the same checks as the notifications above, so a problem appears there for exactly the reason it would be notified, but it differs
|
||||
in three ways:
|
||||
|
||||
- It ignores the "Notify on" toggles. Turning a notification off doesn't hide the problem from the page.
|
||||
- Problems under a **maintenance window** are kept on the list, marked *silenced* and counted separately, instead of being dropped.
|
||||
- It also lists things nothing notifies about: Uptime Kuma monitors that are down, osTicket tickets that are overdue, and any integration that's
|
||||
failing its last few calls (before the threshold that triggers an "integration down" notification).
|
||||
|
||||
It runs the checks live when opened (a recent result is reused for a minute), and shows what it couldn't read at the top, so a missing section means
|
||||
"couldn't check" and not "all clear".
|
||||
|
||||
## What does *not* send a notification
|
||||
|
||||
Worth calling out explicitly, since it's easy to assume everything in
|
||||
|
||||
Reference in new issue
Block a user