fix: truthful fleet status in blind shells + agent-health circuit breaker
box fleet status / approvals check misreported every node as STOPPED / CDP-unreachable from sandboxed shells (own PID+net namespaces: pgrep blind, no route to 10.201.x.x, no sudo). Fleet was healthy throughout. - bin/host_evidence.py (new): host watchdog evidence fallback. Recent timer runs (journal -o json, exact UNIT match) with no newer failure line in cdp-relay-watchdog.log / chromebox-watchdog.log (both silent-when-healthy) prove a node is up. def/dev have no watchdog coverage: browser verdict via chromebox-<node>.log freshness (alive-only), CDP verdict unknown. - super-cli.py: effective status/source/evidence per node. Host evidence decides ONLY the fully-blind pattern (both local probes negative); live local signals always win. New UNKNOWN badge, [*] footnote; approvals UNREACHABLE splits into BLIND / OFFLINE(host agrees) / unreachable-evidence-inconclusive, with honest footer. proc_alive/cdp_ok keep local-probe meaning; status/source/evidence are new JSON fields. - approvals.py: host_cdp_ok flag on the unreachable path. - agent-health.sh: restart circuit breaker. 3 consecutive futile restarts (restart leaves agent still failing) opens the circuit: no more kills for 1800s, ALERT to log+journal, half-open probe after cooldown, reset on any success. Stops the def murder loop (57 restarts / 155 API FAILs for an account-layer failure). - tests/test_fleet_status.py (25), tests/test_agent_health.py (6). - CHROMEBOX-RUNBOOK.md: blind-shell status + futile-restart sections. Tests: 98/98 focused green (agent_health + fleet_status + completion + tool_calls). Live-verified: 4 ACTIVE [*] + 2 UNKNOWN.
This commit is contained in:
@@ -190,6 +190,19 @@ recent-launch guard.
|
||||
**Fix:** fixed 2026-10-04 — 4 retries over 60s + skip kill if launched <2 min ago.
|
||||
If you see this pattern again, the guard may need tuning (longer window).
|
||||
|
||||
### Futile-restart loop (agent-health killing a healthy browser)
|
||||
**Symptoms:** `FAIL (api timeout)` + `restarting browser...` + `CRITICAL -
|
||||
still down after restart` repeating every ~10 min for one node while its
|
||||
CDP port stays up (observed 2026-10-06: def, 57 restarts, 155 API FAILs).
|
||||
The API/account layer is broken; restarts cannot fix it, they just murder
|
||||
a working browser.
|
||||
**Fix:** `agent-health.sh` circuit breaker — after 3 consecutive futile
|
||||
restarts the circuit OPENS (alert in log + journal, no more kills) until
|
||||
a 30-min half-open probe or any successful check. Manual reset:
|
||||
`rm /tmp/agent-health-state/circuit-<node> /tmp/agent-health-state/futile-<node>`.
|
||||
Then fix the actual API-layer failure (account session/auth/chat-state),
|
||||
not the browser.
|
||||
|
||||
### Relay on wrong IP
|
||||
**Symptoms:** relay process exists but on the wrong veth IP (e.g., muse's relay
|
||||
on pip's `10.201.87.2` instead of muse's `10.201.35.2`). Port responds on the
|
||||
@@ -212,6 +225,24 @@ in the NetVM repo — check `git status` if it's gone.
|
||||
**Cause:** old node-ups or queue tests. Harmless but confusing.
|
||||
**Fix:** kill them. Only 9410/9420/9430/9440 should be running.
|
||||
|
||||
## Fleet Status From Blind Shells
|
||||
|
||||
`box fleet status` probes live (pgrep + peer-IP CDP). Sandboxed shells
|
||||
(own PID/net namespaces, no sudo, no route to 10.201.x.x) fail both
|
||||
probes for every node. Instead of misreporting STOPPED, fleet status
|
||||
falls back to host watchdog evidence (`bin/host_evidence.py`):
|
||||
|
||||
- Recent watchdog timer runs (journal) with no newer failure line in
|
||||
`cdp-relay-watchdog.log` / `chromebox-watchdog.log` (both are
|
||||
silent-when-healthy) prove the node is up → `ACTIVE [*]`.
|
||||
- `UNKNOWN` means neither live probes nor host evidence could decide
|
||||
(e.g. def/dev have no relay-monitor coverage).
|
||||
- Host evidence never overrides a live local signal, so a fresh outage
|
||||
observed on the host always wins over a minutes-old watchdog run.
|
||||
|
||||
Same rule drives `box approvals check`: `BLIND` = CDP ok on host, this
|
||||
shell cannot reach it; approval queues are unverified, not clear.
|
||||
|
||||
## Key Reference
|
||||
|
||||
**Node bring-up:** `sudo /home/super/Projects/NetVM/bin/netvm-node-up.sh <node>`
|
||||
|
||||
Reference in New Issue
Block a user