fix: truthful fleet status in blind shells + agent-health circuit breaker

box fleet status / approvals check misreported every node as STOPPED /
CDP-unreachable from sandboxed shells (own PID+net namespaces: pgrep
blind, no route to 10.201.x.x, no sudo). Fleet was healthy throughout.

- bin/host_evidence.py (new): host watchdog evidence fallback. Recent
  timer runs (journal -o json, exact UNIT match) with no newer failure
  line in cdp-relay-watchdog.log / chromebox-watchdog.log (both
  silent-when-healthy) prove a node is up. def/dev have no watchdog
  coverage: browser verdict via chromebox-<node>.log freshness
  (alive-only), CDP verdict unknown.
- super-cli.py: effective status/source/evidence per node. Host
  evidence decides ONLY the fully-blind pattern (both local probes
  negative); live local signals always win. New UNKNOWN badge, [*]
  footnote; approvals UNREACHABLE splits into BLIND / OFFLINE(host
  agrees) / unreachable-evidence-inconclusive, with honest footer.
  proc_alive/cdp_ok keep local-probe meaning; status/source/evidence
  are new JSON fields.
- approvals.py: host_cdp_ok flag on the unreachable path.
- agent-health.sh: restart circuit breaker. 3 consecutive futile
  restarts (restart leaves agent still failing) opens the circuit:
  no more kills for 1800s, ALERT to log+journal, half-open probe
  after cooldown, reset on any success. Stops the def murder loop
  (57 restarts / 155 API FAILs for an account-layer failure).
- tests/test_fleet_status.py (25), tests/test_agent_health.py (6).
- CHROMEBOX-RUNBOOK.md: blind-shell status + futile-restart sections.

Tests: 98/98 focused green (agent_health + fleet_status +
completion + tool_calls). Live-verified: 4 ACTIVE [*] + 2 UNKNOWN.
This commit is contained in:
Muse Sidechat
2026-10-06 18:16:14 +00:00
parent a9f014f9fa
commit c9143a558b
7 changed files with 901 additions and 9 deletions
+14
View File
@@ -616,11 +616,25 @@ def inspect_node_approvals(node: str) -> dict:
"ws_url": "",
"key_request": key_req,
}
# Local CDP probe failed. Ask host evidence whether the node is
# really down or this shell is just blind (sandboxed netns).
host_ok = None
try:
import host_evidence
ev = host_evidence.collect([node]).get(node) or {}
bv, cv = ev.get("browser"), ev.get("cdp")
if cv == "down" or bv == "down":
host_ok = False
elif cv == "healthy" or bv == "healthy":
host_ok = True
except Exception:
host_ok = None
return {
"node": node,
"status": "UNREACHABLE",
"error": str(e),
"has_pending": False,
"host_cdp_ok": host_ok,
}
all_input_waits = []