box fleet status / approvals check misreported every node as STOPPED /
CDP-unreachable from sandboxed shells (own PID+net namespaces: pgrep
blind, no route to 10.201.x.x, no sudo). Fleet was healthy throughout.
- bin/host_evidence.py (new): host watchdog evidence fallback. Recent
timer runs (journal -o json, exact UNIT match) with no newer failure
line in cdp-relay-watchdog.log / chromebox-watchdog.log (both
silent-when-healthy) prove a node is up. def/dev have no watchdog
coverage: browser verdict via chromebox-<node>.log freshness
(alive-only), CDP verdict unknown.
- super-cli.py: effective status/source/evidence per node. Host
evidence decides ONLY the fully-blind pattern (both local probes
negative); live local signals always win. New UNKNOWN badge, [*]
footnote; approvals UNREACHABLE splits into BLIND / OFFLINE(host
agrees) / unreachable-evidence-inconclusive, with honest footer.
proc_alive/cdp_ok keep local-probe meaning; status/source/evidence
are new JSON fields.
- approvals.py: host_cdp_ok flag on the unreachable path.
- agent-health.sh: restart circuit breaker. 3 consecutive futile
restarts (restart leaves agent still failing) opens the circuit:
no more kills for 1800s, ALERT to log+journal, half-open probe
after cooldown, reset on any success. Stops the def murder loop
(57 restarts / 155 API FAILs for an account-layer failure).
- tests/test_fleet_status.py (25), tests/test_agent_health.py (6).
- CHROMEBOX-RUNBOOK.md: blind-shell status + futile-restart sections.
Tests: 98/98 focused green (agent_health + fleet_status +
completion + tool_calls). Live-verified: 4 ACTIVE [*] + 2 UNKNOWN.
- Require 2 CONSECUTIVE muse-chat-api.py API failures before kill -9
(per-node counter in /tmp/agent-health-state, reset on success).
A single 30s API timeout killed 646s healthy browser at 20:48:44 UTC
while its CDP port was still listening.
- Extend post-restart re-check grace to ~60s (15s internal + 45s), matching
chromebox-watchdog.shs proven 60s retry window.
- Add recent_relaunch() guard (mirrors chromebox-watchdog.sh idiom):
skip the kill path when the main browser process launched <2 min ago,
so the two watchdogs can not kill each others fresh browsers.
- check_agent now returns 0/1/2 (healthy/api-fail/port-down); CDP-port
failure still kills immediately. warp-$node checks untouched.
Session: sidechat/chromebox-fixes
agent-health.sh used bare node names for 'ip netns exec' but netns are
named warp-<node> since the NetVM layout; the 6189793 CDP-liveness check
always failed ('No such file or directory'), logging false CRITICALs and
kill -9'ing healthy browsers every 5 min. Use warp-$node.
fleet-alert-check.sh: new 5-min critical-condition detector (per-node CDP
liveness via warp-<node> netns). Consecutive-failure state machine:
page after 2 consecutive failures, re-page every 30 min while critical,
quiet-hours-aware (first alert always pages). Emits ALERT/RECOVERY
records to ~/.local/share/fleet-alert/outbox.jsonl for the container
fleet-alert-relay hook; best-effort box-ctl notify to healthy agents.
Session: sidechat/critical-alerting-pipeline
Root cause: agent-health.service runs Type=oneshot with the default
KillMode=control-group. restart_browser() spawned the replacement
chromium with nohup under the service, so systemd SIGKILLed it the
moment the service exited. Every 5-min tick: FAIL -> restart ->
RECOVERED -> SIGKILL at teardown. opm and pip were permanently dark
and dm.py read masked it (empty output, exit 0).
Fix: launch replacements via systemd-run --user --scope (backgrounded)
so the browser lives in a transient scope outside the service cgroup
and survives teardown. setsid does NOT escape either. Note:
systemd-run --scope waits for the scope even with --no-block
(verified 2026-10-03), hence the backgrounding. Same fix in
chromebox-watchdog.sh (timers currently off).
dm.py: dm_read() now uses run_full() and prints a WARNING to stderr
with rc + last error line instead of failing silently on empty reads.
Trailers: Session: sidechat/opm-blind-fix
New bin/netvm-registry.py parses NODES.md (node -> cdp_port); every
consumer reads from it instead of hardcoding:
- onboard-driver.py: CDP_PORTS dict -> registry lookup (new nodes work)
- muse-signin.py: hardcoded 9410 -> --node/--cdp-port args
- muse-chat-api.py: hardcoded ACCOUNTS -> registry-built
- agent-health.sh: hardcoded muse/pip blocks -> loop over all active
nodes (646 and opm now get health coverage too)
NODES.md: record opm node (9440).