Fix agent silently failing to report on hosts without a cpuinfo model name
report-tasks.sh's new hardware-collection code ran under set -euo pipefail, so a single failing command inside it aborted the whole script before the report was ever sent — with no error message, since nothing in that path had explicit error handling. On real hardware this hit immediately: `grep -m1 "model name" /proc/cpuinfo` exits 1 when there's no match, and many ARM boards (e.g. Raspberry Pi) have no such line at all. Confirmed via journalctl showing the systemd service failing every 15 minutes with exit 1 and zero output, and via a minimal repro of the exact bash control flow. Wrap the hardware/network collection call so any failure inside it is non-fatal: task reporting (the actual core function) must never be taken down by a quirk in the best-effort hardware-gathering code, on this host or any other. Also fixed a related gap found while testing the fallback path: the server only accepted the system field being absent, not explicitly null (what the script now sends if collection fails outright), which would have turned graceful degradation into a rejected report. Verified end-to-end against the real dev server: system:null, system omitted, and a normal populated report all now return 202 and persist correctly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
1 parent
130212baec
commit
ae41a02864
3 files changed
+18
-3
No files matched your search
@@ -191,7 +191,10 @@ collect_system_info() {
|
||||
fi
|
||||
|
||||
local cpu_model="" cpu_cores=0 cpu_load_percent="null"
|
||||
cpu_model=$(grep -m1 "model name" /proc/cpuinfo 2>/dev/null | cut -d: -f2- | sed 's/^ *//')
|
||||
# Some ARM kernels (e.g. Raspberry Pi) have no "model name" line in
|
||||
# /proc/cpuinfo at all — grep then finds no match and exits 1, which is not
|
||||
# a real error here, just an absent field.
|
||||
cpu_model=$(grep -m1 "model name" /proc/cpuinfo 2>/dev/null | cut -d: -f2- | sed 's/^ *//' || true)
|
||||
cpu_cores=$(nproc 2>/dev/null || echo 0)
|
||||
if [[ -r /proc/loadavg && "$cpu_cores" -gt 0 ]]; then
|
||||
local load1
|
||||
@@ -233,7 +236,19 @@ collect_system_info() {
|
||||
|
||||
collect_cron
|
||||
collect_systemd_timers
|
||||
|
||||
# Hardware/network facts are best-effort: a quirk on any given host (e.g. no
|
||||
# "model name" line in /proc/cpuinfo on some ARM boards) must never abort the
|
||||
# whole report over it, so errexit is relaxed just for this call.
|
||||
SYSTEM_JSON="null"
|
||||
set +e
|
||||
collect_system_info
|
||||
system_info_status=$?
|
||||
set -e
|
||||
if [[ "$system_info_status" -ne 0 ]]; then
|
||||
echo "Warning: collecting hardware/network info failed (exit $system_info_status) — reporting tasks without it." >&2
|
||||
SYSTEM_JSON="null"
|
||||
fi
|
||||
|
||||
HOSTNAME_VALUE=$(hostname -f 2>/dev/null || hostname)
|
||||
PAYLOAD=$(jq -n \
|
||||
|
||||
@@ -18,7 +18,7 @@ const reportSchema = z.object({
|
||||
hostname: z.string().max(255).optional(),
|
||||
os_type: z.string().optional(),
|
||||
reported_at: z.string().optional(),
|
||||
system: systemSchema.optional(),
|
||||
system: systemSchema.nullable().optional(),
|
||||
tasks: z.array(
|
||||
z.object({
|
||||
schedule_type: z.enum(["cron", "systemd_timer"]),
|
||||
|
||||
@@ -22,7 +22,7 @@ export interface IncomingSystemInfo {
|
||||
|
||||
export interface AgentReport {
|
||||
hostname?: string;
|
||||
system?: IncomingSystemInfo;
|
||||
system?: IncomingSystemInfo | null;
|
||||
tasks: IncomingTask[];
|
||||
}
|
||||
|
||||
|
||||
Reference in new issue
Block a user