fix: truthful fleet status in blind shells + agent-health circuit breaker
box fleet status / approvals check misreported every node as STOPPED / CDP-unreachable from sandboxed shells (own PID+net namespaces: pgrep blind, no route to 10.201.x.x, no sudo). Fleet was healthy throughout. - bin/host_evidence.py (new): host watchdog evidence fallback. Recent timer runs (journal -o json, exact UNIT match) with no newer failure line in cdp-relay-watchdog.log / chromebox-watchdog.log (both silent-when-healthy) prove a node is up. def/dev have no watchdog coverage: browser verdict via chromebox-<node>.log freshness (alive-only), CDP verdict unknown. - super-cli.py: effective status/source/evidence per node. Host evidence decides ONLY the fully-blind pattern (both local probes negative); live local signals always win. New UNKNOWN badge, [*] footnote; approvals UNREACHABLE splits into BLIND / OFFLINE(host agrees) / unreachable-evidence-inconclusive, with honest footer. proc_alive/cdp_ok keep local-probe meaning; status/source/evidence are new JSON fields. - approvals.py: host_cdp_ok flag on the unreachable path. - agent-health.sh: restart circuit breaker. 3 consecutive futile restarts (restart leaves agent still failing) opens the circuit: no more kills for 1800s, ALERT to log+journal, half-open probe after cooldown, reset on any success. Stops the def murder loop (57 restarts / 155 API FAILs for an account-layer failure). - tests/test_fleet_status.py (25), tests/test_agent_health.py (6). - CHROMEBOX-RUNBOOK.md: blind-shell status + futile-restart sections. Tests: 98/98 focused green (agent_health + fleet_status + completion + tool_calls). Live-verified: 4 ACTIVE [*] + 2 UNKNOWN.
This commit is contained in:
+61
-3
@@ -115,6 +115,57 @@ restart_browser() {
|
||||
echo "$(date -Iseconds) $agent: browser restarted" >> "$LOG"
|
||||
}
|
||||
|
||||
# 2026-10-06: restart circuit breaker. A restart that leaves the agent
|
||||
# still failing is FUTILE (observed 2026-10-06: def's API check failed
|
||||
# 155x while its CDP port was up; 57 kill+restart cycles murdered a
|
||||
# healthy browser for an account-layer failure restarts cannot fix).
|
||||
# After FUTILE_THRESHOLD consecutive futile restarts, open the circuit:
|
||||
# stop killing/restarting and alert, until CIRCUIT_COOLDOWN seconds pass
|
||||
# (one half-open probe restart) or any check succeeds. Manual reset:
|
||||
# rm /tmp/agent-health-state/circuit-<agent> /tmp/agent-health-state/futile-<agent>
|
||||
FUTILE_THRESHOLD=3
|
||||
CIRCUIT_COOLDOWN=1800
|
||||
|
||||
# circuit_allows <agent>: return 0 if a restart may proceed now.
|
||||
circuit_allows() {
|
||||
local agent=$1 now opened retry_in
|
||||
local cf="$STATE_DIR/circuit-$agent"
|
||||
[ -f "$cf" ] || return 0
|
||||
opened=$(cat "$cf" 2>/dev/null || echo 0)
|
||||
now=$(date +%s)
|
||||
if [ $(( now - opened )) -ge $CIRCUIT_COOLDOWN ]; then
|
||||
echo "$(date -Iseconds) $agent: circuit half-open after ${CIRCUIT_COOLDOWN}s cooldown, one probe restart" >> "$LOG"
|
||||
return 0
|
||||
fi
|
||||
retry_in=$(( (opened + CIRCUIT_COOLDOWN - now + 59) / 60 ))
|
||||
echo "$(date -Iseconds) $agent: CIRCUIT OPEN - skipping kill/restart (restarts futile, probe retry in ~${retry_in}m; manual reset: rm $cf)" >> "$LOG"
|
||||
return 1
|
||||
}
|
||||
|
||||
# circuit_note_restart <agent> <ok|fail>: record a restart outcome.
|
||||
circuit_note_restart() {
|
||||
local agent=$1 outcome=$2 count=0
|
||||
local ff="$STATE_DIR/futile-$agent" cf="$STATE_DIR/circuit-$agent"
|
||||
if [ "$outcome" = "ok" ]; then
|
||||
rm -f "$ff" "$cf" 2>/dev/null
|
||||
return 0
|
||||
fi
|
||||
[ -f "$ff" ] && count=$(cat "$ff" 2>/dev/null || echo 0)
|
||||
count=$(( count + 1 ))
|
||||
echo "$count" > "$ff"
|
||||
if [ "$count" -ge "$FUTILE_THRESHOLD" ]; then
|
||||
date +%s > "$cf"
|
||||
local msg="$agent: ALERT - $count consecutive futile restarts, circuit OPEN for ${CIRCUIT_COOLDOWN}s (no more kills until probe; manual reset: rm $cf $ff)"
|
||||
echo "$(date -Iseconds) $msg" >> "$LOG"
|
||||
echo "agent-health ALERT: $msg"
|
||||
fi
|
||||
}
|
||||
|
||||
# Allow sourcing for tests without running checks.
|
||||
if [ "${AGENT_HEALTH_LIB_ONLY:-}" = "1" ]; then
|
||||
return 0 2>/dev/null || exit 0
|
||||
fi
|
||||
|
||||
# Main
|
||||
echo "=== Health check $(date -Iseconds) ===" >> "$LOG"
|
||||
|
||||
@@ -126,8 +177,9 @@ check_one() {
|
||||
local rc=$?
|
||||
|
||||
if [ $rc -eq 0 ]; then
|
||||
# Healthy: reset the consecutive-API-failure counter.
|
||||
rm -f "$STATE_DIR/failcount-$agent" 2>/dev/null
|
||||
# Healthy: reset the consecutive-API-failure counter and close
|
||||
# any open restart circuit.
|
||||
rm -f "$STATE_DIR/failcount-$agent" "$STATE_DIR/futile-$agent" "$STATE_DIR/circuit-$agent" 2>/dev/null
|
||||
return 0
|
||||
fi
|
||||
|
||||
@@ -158,6 +210,11 @@ check_one() {
|
||||
return 0
|
||||
fi
|
||||
|
||||
# Circuit breaker: repeated futile restarts stop here until cooldown.
|
||||
if ! circuit_allows "$agent"; then
|
||||
return 0
|
||||
fi
|
||||
|
||||
restart_browser "$agent" "$cdp_port"
|
||||
# 2026-10-04: post-restart re-check grace extended to ~60s total
|
||||
# (restart_browser sleeps 15s internally + 45s here), matching
|
||||
@@ -166,9 +223,10 @@ check_one() {
|
||||
sleep 45
|
||||
if ! check_agent "$agent" "$agent" "$cdp_port"; then
|
||||
echo "$(date -Iseconds) $agent: CRITICAL - still down after restart" >> "$LOG"
|
||||
# TODO: Alert operator (e.g., via board post or email)
|
||||
circuit_note_restart "$agent" fail
|
||||
else
|
||||
echo "$(date -Iseconds) $agent: RECOVERED after restart" >> "$LOG"
|
||||
circuit_note_restart "$agent" ok
|
||||
fi
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user