fix: truthful fleet status in blind shells + agent-health circuit breaker

box fleet status / approvals check misreported every node as STOPPED /
CDP-unreachable from sandboxed shells (own PID+net namespaces: pgrep
blind, no route to 10.201.x.x, no sudo). Fleet was healthy throughout.

- bin/host_evidence.py (new): host watchdog evidence fallback. Recent
  timer runs (journal -o json, exact UNIT match) with no newer failure
  line in cdp-relay-watchdog.log / chromebox-watchdog.log (both
  silent-when-healthy) prove a node is up. def/dev have no watchdog
  coverage: browser verdict via chromebox-<node>.log freshness
  (alive-only), CDP verdict unknown.
- super-cli.py: effective status/source/evidence per node. Host
  evidence decides ONLY the fully-blind pattern (both local probes
  negative); live local signals always win. New UNKNOWN badge, [*]
  footnote; approvals UNREACHABLE splits into BLIND / OFFLINE(host
  agrees) / unreachable-evidence-inconclusive, with honest footer.
  proc_alive/cdp_ok keep local-probe meaning; status/source/evidence
  are new JSON fields.
- approvals.py: host_cdp_ok flag on the unreachable path.
- agent-health.sh: restart circuit breaker. 3 consecutive futile
  restarts (restart leaves agent still failing) opens the circuit:
  no more kills for 1800s, ALERT to log+journal, half-open probe
  after cooldown, reset on any success. Stops the def murder loop
  (57 restarts / 155 API FAILs for an account-layer failure).
- tests/test_fleet_status.py (25), tests/test_agent_health.py (6).
- CHROMEBOX-RUNBOOK.md: blind-shell status + futile-restart sections.

Tests: 98/98 focused green (agent_health + fleet_status +
completion + tool_calls). Live-verified: 4 ACTIVE [*] + 2 UNKNOWN.
This commit is contained in:
Muse Sidechat
2026-10-06 18:16:14 +00:00
parent a9f014f9fa
commit c9143a558b
7 changed files with 901 additions and 9 deletions
+61 -3
View File
@@ -115,6 +115,57 @@ restart_browser() {
echo "$(date -Iseconds) $agent: browser restarted" >> "$LOG"
}
# 2026-10-06: restart circuit breaker. A restart that leaves the agent
# still failing is FUTILE (observed 2026-10-06: def's API check failed
# 155x while its CDP port was up; 57 kill+restart cycles murdered a
# healthy browser for an account-layer failure restarts cannot fix).
# After FUTILE_THRESHOLD consecutive futile restarts, open the circuit:
# stop killing/restarting and alert, until CIRCUIT_COOLDOWN seconds pass
# (one half-open probe restart) or any check succeeds. Manual reset:
# rm /tmp/agent-health-state/circuit-<agent> /tmp/agent-health-state/futile-<agent>
FUTILE_THRESHOLD=3
CIRCUIT_COOLDOWN=1800
# circuit_allows <agent>: return 0 if a restart may proceed now.
circuit_allows() {
local agent=$1 now opened retry_in
local cf="$STATE_DIR/circuit-$agent"
[ -f "$cf" ] || return 0
opened=$(cat "$cf" 2>/dev/null || echo 0)
now=$(date +%s)
if [ $(( now - opened )) -ge $CIRCUIT_COOLDOWN ]; then
echo "$(date -Iseconds) $agent: circuit half-open after ${CIRCUIT_COOLDOWN}s cooldown, one probe restart" >> "$LOG"
return 0
fi
retry_in=$(( (opened + CIRCUIT_COOLDOWN - now + 59) / 60 ))
echo "$(date -Iseconds) $agent: CIRCUIT OPEN - skipping kill/restart (restarts futile, probe retry in ~${retry_in}m; manual reset: rm $cf)" >> "$LOG"
return 1
}
# circuit_note_restart <agent> <ok|fail>: record a restart outcome.
circuit_note_restart() {
local agent=$1 outcome=$2 count=0
local ff="$STATE_DIR/futile-$agent" cf="$STATE_DIR/circuit-$agent"
if [ "$outcome" = "ok" ]; then
rm -f "$ff" "$cf" 2>/dev/null
return 0
fi
[ -f "$ff" ] && count=$(cat "$ff" 2>/dev/null || echo 0)
count=$(( count + 1 ))
echo "$count" > "$ff"
if [ "$count" -ge "$FUTILE_THRESHOLD" ]; then
date +%s > "$cf"
local msg="$agent: ALERT - $count consecutive futile restarts, circuit OPEN for ${CIRCUIT_COOLDOWN}s (no more kills until probe; manual reset: rm $cf $ff)"
echo "$(date -Iseconds) $msg" >> "$LOG"
echo "agent-health ALERT: $msg"
fi
}
# Allow sourcing for tests without running checks.
if [ "${AGENT_HEALTH_LIB_ONLY:-}" = "1" ]; then
return 0 2>/dev/null || exit 0
fi
# Main
echo "=== Health check $(date -Iseconds) ===" >> "$LOG"
@@ -126,8 +177,9 @@ check_one() {
local rc=$?
if [ $rc -eq 0 ]; then
# Healthy: reset the consecutive-API-failure counter.
rm -f "$STATE_DIR/failcount-$agent" 2>/dev/null
# Healthy: reset the consecutive-API-failure counter and close
# any open restart circuit.
rm -f "$STATE_DIR/failcount-$agent" "$STATE_DIR/futile-$agent" "$STATE_DIR/circuit-$agent" 2>/dev/null
return 0
fi
@@ -158,6 +210,11 @@ check_one() {
return 0
fi
# Circuit breaker: repeated futile restarts stop here until cooldown.
if ! circuit_allows "$agent"; then
return 0
fi
restart_browser "$agent" "$cdp_port"
# 2026-10-04: post-restart re-check grace extended to ~60s total
# (restart_browser sleeps 15s internally + 45s here), matching
@@ -166,9 +223,10 @@ check_one() {
sleep 45
if ! check_agent "$agent" "$agent" "$cdp_port"; then
echo "$(date -Iseconds) $agent: CRITICAL - still down after restart" >> "$LOG"
# TODO: Alert operator (e.g., via board post or email)
circuit_note_restart "$agent" fail
else
echo "$(date -Iseconds) $agent: RECOVERED after restart" >> "$LOG"
circuit_note_restart "$agent" ok
fi
}
+14
View File
@@ -616,11 +616,25 @@ def inspect_node_approvals(node: str) -> dict:
"ws_url": "",
"key_request": key_req,
}
# Local CDP probe failed. Ask host evidence whether the node is
# really down or this shell is just blind (sandboxed netns).
host_ok = None
try:
import host_evidence
ev = host_evidence.collect([node]).get(node) or {}
bv, cv = ev.get("browser"), ev.get("cdp")
if cv == "down" or bv == "down":
host_ok = False
elif cv == "healthy" or bv == "healthy":
host_ok = True
except Exception:
host_ok = None
return {
"node": node,
"status": "UNREACHABLE",
"error": str(e),
"has_pending": False,
"host_cdp_ok": host_ok,
}
all_input_waits = []
+348
View File
@@ -0,0 +1,348 @@
#!/usr/bin/env python3
"""Host-side fleet evidence for network/PID-blind shells.
`box fleet status` probes each node live (pgrep for the chromium process,
HTTP to the CDP relay on the peer IP). Both probes assume the caller's
network + PID namespace is the bl host's. From a sandboxed shell (own PID
and net namespaces, no sudo, no route to 10.201.x.x) both probes always
fail, so every node misreports as STOPPED even with a healthy fleet.
This module provides the fallback signal: evidence written by the
host-side watchdogs that run on bl unsandboxed via systemd timers:
- cdp-relay-watchdog (every 5 min, all active registry nodes):
log ``cdp-relay-watchdog.log`` + journal unit
``cdp-relay-watchdog.service``. Proves the CDP relay path end to end.
- chromebox-watchdog (every 2 min, same nodes, one timer per node):
log ``chromebox-watchdog.log`` + journal units
``chromebox-watchdog@<node>.service``. Curls CDP inside the node netns,
so it proves browser + in-netns CDP.
- ``chromebox-<node>.log`` mtime: chromium's own stdout. Fresh output
proves the browser process is alive. Used only for nodes outside
watchdog coverage — and only to conclude "alive", never "dead".
Both watchdogs are silent-when-healthy: a failure is ALWAYS logged, so a
recent timer run (journal "Starting" line) with no newer failure line for
the node means that run found the node healthy.
Verdicts: "healthy" | "degraded" | "down" | "unknown".
"""
import json
import os
import re
import subprocess
import time
from datetime import datetime, timezone
from pathlib import Path
NETVM_ROOT = Path("/home/super/Projects/NetVM")
RELAY_LOG = NETVM_ROOT / "cdp-relay-watchdog.log"
CHROMEBOX_LOG = NETVM_ROOT / "chromebox-watchdog.log"
ALL_NODES = ("muse", "pip", "646", "opm", "def", "dev")
def _covered_nodes():
"""Nodes with watchdog coverage, from the fleet registry.
Both watchdogs supervise every active registry node. Falls back to
ALL_NODES when the registry is unreadable, so a broken registry can
never silently narrow fleet status to a subset of the fleet.
"""
try:
import importlib.util
spec = importlib.util.spec_from_file_location(
"netvm_registry", NETVM_ROOT / "bin" / "netvm-registry.py")
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return tuple(sorted(mod.active_nodes()))
except Exception:
return ALL_NODES
RELAY_NODES = _covered_nodes()
CHROMEBOX_NODES = _covered_nodes()
RELAY_UNIT = "cdp-relay-watchdog.service"
CHROMEBOX_UNIT_TMPL = "chromebox-watchdog@{node}.service"
# A watchdog run older than this proves nothing (timer may be dead).
RELAY_STALE_MIN = 15
CHROMEBOX_STALE_MIN = 8
# Chromium stdout older than this proves nothing (idle browsers go quiet).
CHROME_LOG_FRESH_MIN = 20
_LOG_TS_RE = re.compile(r"^\[(\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2})Z\]")
def _utcnow():
return datetime.now(timezone.utc)
def parse_log_ts(line):
"""Parse a ``[YYYY-MM-DDTHH:MM:SSZ]`` log prefix. None if absent."""
m = _LOG_TS_RE.match(line)
if not m:
return None
try:
return datetime.strptime(m.group(1), "%Y-%m-%dT%H:%M:%S").replace(
tzinfo=timezone.utc)
except ValueError:
return None
def classify_chromebox_line(line):
"""Classify one chromebox-watchdog log line.
Returns "healthy" | "degraded" | "down", or None when the line
carries no verdict (rotation markers, relay stdout passthrough).
"""
if "log rotated" in line:
return None
if "relaunch FAILED" in line:
return "down"
if "relaunch OK" in line:
return "healthy"
if "recovered" in line and "relaunch not needed" in line:
return "healthy"
if "not healthy yet" in line:
return "degraded"
if "proceeding with chrome relaunch" in line:
return "degraded"
if "relaunching chromebox" in line:
return "degraded"
if "warp partition detected" in line:
return "degraded"
if "skipping relaunch (probably still starting)" in line:
return "degraded"
return None
def classify_relay_line(line):
"""Classify one cdp-relay-watchdog log line (None = no verdict)."""
if "relay restart FAILED" in line:
return "down"
if "FAIL_LOUD" in line:
return "down"
if "relay restarted OK" in line:
return "healthy"
if "relay unhealthy" in line and "restarting" in line:
# Always followed by an OK/FAILED line; a trailing one means the
# restart crashed mid-flight.
return "down"
return None
def last_verdict(lines, node, classify):
"""Newest (verdict, ts, line) for ``[node]``. None if no verdict line."""
tag = "[%s]" % node
best = None
for line in lines:
if tag not in line:
continue
verdict = classify(line)
if verdict is None:
continue
ts = parse_log_ts(line)
if ts is None:
continue
if best is None or ts >= best[1]:
best = (verdict, ts, line.strip()[:160])
return best
def _tail_lines(path, max_bytes=65536):
try:
size = os.path.getsize(path)
with open(path, "rb") as f:
if size > max_bytes:
f.seek(size - max_bytes)
f.readline() # drop partial first line
return f.read().decode("utf-8", errors="replace").splitlines()
except OSError:
return []
def query_journal_starts(units, since_min=25, timeout=20):
"""Map each unit -> newest run-start (aware UTC). Missing on failure.
Uses ``-o json``: the short-format "Starting" line carries the unit
description, not the unit name, so exact per-unit matching needs the
structured UNIT field.
"""
cmd = ["journalctl", "--no-pager", "-o", "json",
"--since", "%d min ago" % since_min]
for u in units:
cmd.extend(["-u", u])
try:
r = subprocess.run(cmd, capture_output=True, text=True, timeout=timeout)
except (OSError, subprocess.TimeoutExpired):
return {}
if r.returncode != 0:
return {}
want = set(units)
starts = {}
for line in (r.stdout or "").splitlines():
try:
e = json.loads(line)
except ValueError:
continue
if e.get("UNIT") not in want:
continue
if not (e.get("MESSAGE") or "").startswith("Starting"):
continue
try:
ts = datetime.fromtimestamp(
int(e["__REALTIME_TIMESTAMP"]) / 1e6, tz=timezone.utc)
except (KeyError, ValueError, TypeError, OverflowError):
continue
u = e["UNIT"]
if u not in starts or ts > starts[u]:
starts[u] = ts
return starts
def _verdict_since_run(verdict_row, run_ts):
"""True when the verdict line is newer than (or from) the last run."""
if verdict_row is None or run_ts is None:
return False
return verdict_row[1] >= run_ts
def browser_verdict(node, chromebox_lines, run_ts, chrome_log_mtime=None,
now=None):
"""(verdict, detail) for the node's browser process."""
now = now or _utcnow()
row = last_verdict(chromebox_lines, node, classify_chromebox_line)
if node in CHROMEBOX_NODES:
if run_ts is None:
return ("unknown", "no chromebox-watchdog run in journal window")
if (now - run_ts).total_seconds() > CHROMEBOX_STALE_MIN * 60:
return ("unknown", "chromebox-watchdog run is stale")
if _verdict_since_run(row, run_ts):
return (row[0], "watchdog: %s" % row[2])
return ("healthy", "watchdog run silent (silent-when-healthy)")
# Nodes outside watchdog coverage: chromium stdout proves alive only.
if chrome_log_mtime is not None and (
now - chrome_log_mtime).total_seconds() < CHROME_LOG_FRESH_MIN * 60:
return ("healthy", "chromebox-%s.log fresh" % node)
return ("unknown", "no watchdog coverage for %s" % node)
def cdp_verdict(node, relay_lines, run_ts, now=None):
"""(verdict, detail) for the node's host-reachable CDP relay path."""
now = now or _utcnow()
if node not in RELAY_NODES:
return ("unknown", "no relay-monitor coverage for %s" % node)
if run_ts is None:
return ("unknown", "no cdp-relay-watchdog run in journal window")
if (now - run_ts).total_seconds() > RELAY_STALE_MIN * 60:
return ("unknown", "cdp-relay-watchdog run is stale")
row = last_verdict(relay_lines, node, classify_relay_line)
if _verdict_since_run(row, run_ts):
return (row[0], "relay watchdog: %s" % row[2])
return ("healthy", "relay watchdog run silent (silent-when-healthy)")
def chrome_log_mtime(node):
"""Mtime of chromium's stdout log as aware UTC. None if missing."""
try:
return datetime.fromtimestamp(
os.path.getmtime(NETVM_ROOT / ("chromebox-%s.log" % node)),
tz=timezone.utc)
except OSError:
return None
_CACHE = {"at": 0.0, "nodes": frozenset(), "data": {}}
_CACHE_TTL_S = 60
def collect(nodes=None, _journal_starts=None, _relay_lines=None,
_chromebox_lines=None, _chrome_mtimes=None, _now=None):
"""Per-node host evidence. Underscore args are seams for tests."""
nodes = list(nodes or ALL_NODES)
live = (_journal_starts is None and _relay_lines is None
and _chromebox_lines is None and _chrome_mtimes is None
and _now is None)
if live:
key = frozenset(nodes)
if (key <= _CACHE["nodes"]
and time.monotonic() - _CACHE["at"] < _CACHE_TTL_S):
return {n: _CACHE["data"][n] for n in nodes if n in _CACHE["data"]}
now = _now or _utcnow()
if _journal_starts is None:
units = [RELAY_UNIT] + [CHROMEBOX_UNIT_TMPL.format(node=n)
for n in nodes if n in CHROMEBOX_NODES]
_journal_starts = query_journal_starts(units)
if _relay_lines is None:
_relay_lines = _tail_lines(RELAY_LOG)
if _chromebox_lines is None:
_chromebox_lines = _tail_lines(CHROMEBOX_LOG)
out = {}
for node in nodes:
if _chrome_mtimes is not None and node in _chrome_mtimes:
mtime = _chrome_mtimes[node]
else:
mtime = chrome_log_mtime(node) if node not in CHROMEBOX_NODES else None
b_verd, b_det = browser_verdict(
node, _chromebox_lines,
_journal_starts.get(CHROMEBOX_UNIT_TMPL.format(node=node)),
chrome_log_mtime=mtime, now=now)
c_verd, c_det = cdp_verdict(
node, _relay_lines, _journal_starts.get(RELAY_UNIT), now=now)
out[node] = {
"browser": b_verd,
"browser_detail": b_det,
"cdp": c_verd,
"cdp_detail": c_det,
}
if live:
_CACHE["at"] = time.monotonic()
_CACHE["nodes"] = frozenset(nodes)
_CACHE["data"] = out
return out
def effective_status(local_proc, local_cdp, browser_v, cdp_v):
"""Map (local probes, host verdicts) -> (status, source).
Host evidence only ever overrides the fully-blind pattern (both
local probes negative — the sandbox signature). It never overrides
a live local signal, so a fresh outage on the host always wins.
"""
if local_cdp:
# CDP answers: the browser is definitionally alive.
return ("ACTIVE", "local")
if local_proc:
return ("CDP_DOWN", "local")
# Both local probes negative: consult host evidence.
if browser_v == "down":
return ("STOPPED", "host-evidence")
if browser_v == "unknown":
return ("UNKNOWN", "host-evidence")
if cdp_v == "healthy":
return ("ACTIVE", "host-evidence")
if cdp_v == "down":
return ("CDP_DOWN", "host-evidence")
return ("UNKNOWN", "host-evidence")
def main(argv=None):
import argparse
ap = argparse.ArgumentParser(description="Show host-side fleet evidence")
ap.add_argument("--json", action="store_true")
args = ap.parse_args(argv)
data = collect()
if args.json:
print(json.dumps({"ok": True, "evidence": data}, indent=2))
return
for node, ev in data.items():
print("%-6s browser=%-8s cdp=%-8s" % (node, ev["browser"], ev["cdp"]))
print(" browser: %s" % ev["browser_detail"])
print(" cdp: %s" % ev["cdp_detail"])
if __name__ == "__main__":
main()
+80 -6
View File
@@ -252,6 +252,53 @@ def get_queue_depth(node: str) -> int:
# ---------------------------------------------------------------------------
# Domain: FLEET
# ---------------------------------------------------------------------------
def _host_evidence_for(results):
"""Host watchdog evidence, or None when unneeded/unavailable.
Queried only when at least one node is fully blind (both local
probes negative). Never raises: evidence must not break status.
"""
if not any(not r["proc_alive"] and not r["cdp_ok"] for r in results):
return None
try:
import host_evidence
return host_evidence.collect([r["node"] for r in results])
except Exception:
return None
def _resolve_statuses(results) -> None:
"""Attach effective status/source/evidence to each collected node.
proc_alive/cdp_ok keep their local-probe meaning; status is the
display verdict. Host evidence only decides the fully-blind pattern
(both local probes negative); live local signals always win.
"""
try:
import host_evidence
have_mod = True
except ImportError:
have_mod = False
ev_map = _host_evidence_for(results) if have_mod else None
for r in results:
ev = (ev_map or {}).get(r["node"]) if ev_map else None
if not r["proc_alive"] and r["cdp_ok"]:
# CDP answers: the browser is definitionally alive even when
# pgrep cannot see it (PID-blind shell, renamed profile path).
r["status"], r["source"] = "ACTIVE", "local"
elif r["proc_alive"] and not r["cdp_ok"]:
r["status"], r["source"] = "CDP_DOWN", "local"
elif r["proc_alive"] and r["cdp_ok"]:
r["status"], r["source"] = "ACTIVE", "local"
elif ev is None:
r["status"], r["source"] = "STOPPED", "local"
else:
r["status"], r["source"] = host_evidence.effective_status(
False, False, ev.get("browser", "unknown"),
ev.get("cdp", "unknown"))
r["evidence"] = ev
def collect_fleet_data() -> list:
results = []
for node in VALID_NODES:
@@ -287,6 +334,7 @@ def collect_fleet_data() -> list:
"approval_pending": approval_pending,
"approval_detail": approval_detail,
})
_resolve_statuses(results)
return results
def cmd_fleet_status(args):
@@ -301,14 +349,16 @@ def cmd_fleet_status(args):
rows = []
has_any_approval = False
for item in data:
# Status calculation
# Status calculation (effective status; see _resolve_statuses)
if item.get("approval_pending"):
status = badge_warn("APPROVAL_REQ")
has_any_approval = True
elif item["proc_alive"] and item["cdp_ok"]:
elif item.get("status") == "ACTIVE":
status = badge_ok("ACTIVE")
elif item["proc_alive"] and not item["cdp_ok"]:
elif item.get("status") == "CDP_DOWN":
status = badge_warn("CDP_DOWN")
elif item.get("status") == "UNKNOWN":
status = badge_dim("UNKNOWN")
else:
status = badge_err("STOPPED")
@@ -338,6 +388,8 @@ def cmd_fleet_status(args):
])
print_table(headers, rows)
if any(r.get("source") == "host-evidence" for r in data):
print(c_dim(" [*] status via host watchdog evidence (local probes blind in this shell)"))
if has_any_approval:
print("\n" + c_yellow(" ⚠ Agent(s) held up on browser approval. Run 'box approvals' to inspect/allow."))
print("\n" + c_dim(" Commands: super fleet watch | super fleet restart <node> | box approvals [check|allow|auto]") + "\n")
@@ -446,9 +498,16 @@ def cmd_approvals(args):
btns = c_dim("answer in task")
input_wait_count += 1
elif st == "UNREACHABLE":
badge = badge_dim("OFFLINE")
if it.get("host_cdp_ok") is True:
badge = badge_dim("BLIND")
purp = c_dim("CDP ok on host; blind here")
elif it.get("host_cdp_ok") is False:
badge = badge_err("OFFLINE")
purp = c_dim("CDP down (host agrees)")
else:
badge = badge_dim("OFFLINE")
purp = c_dim("CDP unreachable")
target = "-"
purp = c_dim("CDP unreachable")
trust = "-"
btns = "-"
elif st == "ERROR":
@@ -478,7 +537,22 @@ def cmd_approvals(args):
print(c_cyan(" All pending approvals are for trusted infrastructure. Run 'box approvals auto' to resolve."))
print()
elif input_wait_count == 0:
print("\n" + c_green(" ✔ All agent approval queues clear. No agents blocked.") + "\n")
unreach = [it for it in fleet if it.get("status") == "UNREACHABLE"]
blind_ok = [it["node"] for it in unreach
if it.get("host_cdp_ok") is True]
blind_unknown = [it["node"] for it in unreach
if it.get("host_cdp_ok") is None]
if blind_ok:
print("\n" + c_dim(" ? %d node(s) blind from this shell "
"(CDP ok on host); queues unverified: %s."
% (len(blind_ok), ", ".join(blind_ok))) + "\n")
if blind_unknown:
print("\n" + c_dim(" ? %d node(s) unreachable, host evidence "
"inconclusive: %s."
% (len(blind_unknown),
", ".join(blind_unknown))) + "\n")
if not blind_ok and not blind_unknown:
print("\n" + c_green(" ✔ All agent approval queues clear. No agents blocked.") + "\n")
# Verbose or inspect breakdown
is_verbose = getattr(args, "verbose", False) or action == "inspect"