fleet alerting: fix netns naming in agent-health.sh + add fleet-alert-check.sh

agent-health.sh used bare node names for 'ip netns exec' but netns are

named warp-<node> since the NetVM layout; the 6189793 CDP-liveness check

always failed ('No such file or directory'), logging false CRITICALs and

kill -9'ing healthy browsers every 5 min. Use warp-$node.

fleet-alert-check.sh: new 5-min critical-condition detector (per-node CDP

liveness via warp-<node> netns). Consecutive-failure state machine:

page after 2 consecutive failures, re-page every 30 min while critical,

quiet-hours-aware (first alert always pages). Emits ALERT/RECOVERY

records to ~/.local/share/fleet-alert/outbox.jsonl for the container

fleet-alert-relay hook; best-effort box-ctl notify to healthy agents.

Session: sidechat/critical-alerting-pipeline
This commit is contained in:
operator-main
2026-10-04 20:09:36 +00:00
parent 8bb64c4826
commit 5ab157b3b0
2 changed files with 213 additions and 2 deletions
+2 -2
View File
@@ -21,7 +21,7 @@ check_cdp_port() {
# A browser can be running but not bound to CDP (zombie state).
local node=$1
local cdp_port=$2
if sudo ip netns exec "$node" ss -tln 2>/dev/null | grep -q ":$cdp_port "; then
if sudo ip netns exec "warp-$node" ss -tln 2>/dev/null | grep -q ":$cdp_port "; then
return 0
else
return 1
@@ -61,7 +61,7 @@ restart_browser() {
done
sleep 3
# Verify port is free before restart
if sudo ip netns exec "$agent" ss -tln 2>/dev/null | grep -q ":$cdp_port "; then
if sudo ip netns exec "warp-$agent" ss -tln 2>/dev/null | grep -q ":$cdp_port "; then
echo "$(date -Iseconds) $agent: WARNING - port $cdp_port still bound after kill" >> "$LOG"
fi
# Restart via netvm-chrome.sh in its own systemd scope.