fleet alerting: fix netns naming in agent-health.sh + add fleet-alert-check.sh
agent-health.sh used bare node names for 'ip netns exec' but netns are
named warp-<node> since the NetVM layout; the 6189793 CDP-liveness check
always failed ('No such file or directory'), logging false CRITICALs and
kill -9'ing healthy browsers every 5 min. Use warp-$node.
fleet-alert-check.sh: new 5-min critical-condition detector (per-node CDP
liveness via warp-<node> netns). Consecutive-failure state machine:
page after 2 consecutive failures, re-page every 30 min while critical,
quiet-hours-aware (first alert always pages). Emits ALERT/RECOVERY
records to ~/.local/share/fleet-alert/outbox.jsonl for the container
fleet-alert-relay hook; best-effort box-ctl notify to healthy agents.
Session: sidechat/critical-alerting-pipeline
This commit is contained in:
+2
-2
@@ -21,7 +21,7 @@ check_cdp_port() {
|
||||
# A browser can be running but not bound to CDP (zombie state).
|
||||
local node=$1
|
||||
local cdp_port=$2
|
||||
if sudo ip netns exec "$node" ss -tln 2>/dev/null | grep -q ":$cdp_port "; then
|
||||
if sudo ip netns exec "warp-$node" ss -tln 2>/dev/null | grep -q ":$cdp_port "; then
|
||||
return 0
|
||||
else
|
||||
return 1
|
||||
@@ -61,7 +61,7 @@ restart_browser() {
|
||||
done
|
||||
sleep 3
|
||||
# Verify port is free before restart
|
||||
if sudo ip netns exec "$agent" ss -tln 2>/dev/null | grep -q ":$cdp_port "; then
|
||||
if sudo ip netns exec "warp-$agent" ss -tln 2>/dev/null | grep -q ":$cdp_port "; then
|
||||
echo "$(date -Iseconds) $agent: WARNING - port $cdp_port still bound after kill" >> "$LOG"
|
||||
fi
|
||||
# Restart via netvm-chrome.sh in its own systemd scope.
|
||||
|
||||
Reference in New Issue
Block a user