8 Commits

Author SHA1 Message Date
Muse Sidechat c9143a558b fix: truthful fleet status in blind shells + agent-health circuit breaker
box fleet status / approvals check misreported every node as STOPPED /
CDP-unreachable from sandboxed shells (own PID+net namespaces: pgrep
blind, no route to 10.201.x.x, no sudo). Fleet was healthy throughout.

- bin/host_evidence.py (new): host watchdog evidence fallback. Recent
  timer runs (journal -o json, exact UNIT match) with no newer failure
  line in cdp-relay-watchdog.log / chromebox-watchdog.log (both
  silent-when-healthy) prove a node is up. def/dev have no watchdog
  coverage: browser verdict via chromebox-<node>.log freshness
  (alive-only), CDP verdict unknown.
- super-cli.py: effective status/source/evidence per node. Host
  evidence decides ONLY the fully-blind pattern (both local probes
  negative); live local signals always win. New UNKNOWN badge, [*]
  footnote; approvals UNREACHABLE splits into BLIND / OFFLINE(host
  agrees) / unreachable-evidence-inconclusive, with honest footer.
  proc_alive/cdp_ok keep local-probe meaning; status/source/evidence
  are new JSON fields.
- approvals.py: host_cdp_ok flag on the unreachable path.
- agent-health.sh: restart circuit breaker. 3 consecutive futile
  restarts (restart leaves agent still failing) opens the circuit:
  no more kills for 1800s, ALERT to log+journal, half-open probe
  after cooldown, reset on any success. Stops the def murder loop
  (57 restarts / 155 API FAILs for an account-layer failure).
- tests/test_fleet_status.py (25), tests/test_agent_health.py (6).
- CHROMEBOX-RUNBOOK.md: blind-shell status + futile-restart sections.

Tests: 98/98 focused green (agent_health + fleet_status +
completion + tool_calls). Live-verified: 4 ACTIVE [*] + 2 UNKNOWN.
2026-10-06 18:16:14 +00:00
operator-main 76f854da6b Harden agent-health.sh: soften API-timeout kill path, add recent-relaunch guard
- Require 2 CONSECUTIVE muse-chat-api.py API failures before kill -9
  (per-node counter in /tmp/agent-health-state, reset on success).
  A single 30s API timeout killed 646s healthy browser at 20:48:44 UTC
  while its CDP port was still listening.
- Extend post-restart re-check grace to ~60s (15s internal + 45s), matching
  chromebox-watchdog.shs proven 60s retry window.
- Add recent_relaunch() guard (mirrors chromebox-watchdog.sh idiom):
  skip the kill path when the main browser process launched <2 min ago,
  so the two watchdogs can not kill each others fresh browsers.
- check_agent now returns 0/1/2 (healthy/api-fail/port-down); CDP-port
  failure still kills immediately. warp-$node checks untouched.

Session: sidechat/chromebox-fixes
2026-10-04 22:46:44 +00:00
operator-main 5ab157b3b0 fleet alerting: fix netns naming in agent-health.sh + add fleet-alert-check.sh
agent-health.sh used bare node names for 'ip netns exec' but netns are

named warp-<node> since the NetVM layout; the 6189793 CDP-liveness check

always failed ('No such file or directory'), logging false CRITICALs and

kill -9'ing healthy browsers every 5 min. Use warp-$node.

fleet-alert-check.sh: new 5-min critical-condition detector (per-node CDP

liveness via warp-<node> netns). Consecutive-failure state machine:

page after 2 consecutive failures, re-page every 30 min while critical,

quiet-hours-aware (first alert always pages). Emits ALERT/RECOVERY

records to ~/.local/share/fleet-alert/outbox.jsonl for the container

fleet-alert-relay hook; best-effort box-ctl notify to healthy agents.

Session: sidechat/critical-alerting-pipeline
2026-10-04 20:09:36 +00:00
operator-main 6189793748 agent-health.sh: Verify CDP port liveness, not just API success\n\nThe watchdog only checked if the API responded, missing zombie\nbrowsers (process alive but CDP not listening). Now explicitly\nchecks via ss -tln in the netns before trying the API.\nAlso: kill by exact PIDs instead of unreliable pkill patterns,\nand warn if port still bound after kill. 2026-10-04 02:23:20 +00:00
operator-main f19aff9b79 Fix supervisor SIGKILL loop killing restarted browsers (opm/pip)
Root cause: agent-health.service runs Type=oneshot with the default
KillMode=control-group. restart_browser() spawned the replacement
chromium with nohup under the service, so systemd SIGKILLed it the
moment the service exited. Every 5-min tick: FAIL -> restart ->
RECOVERED -> SIGKILL at teardown. opm and pip were permanently dark
and dm.py read masked it (empty output, exit 0).

Fix: launch replacements via systemd-run --user --scope (backgrounded)
so the browser lives in a transient scope outside the service cgroup
and survives teardown. setsid does NOT escape either. Note:
systemd-run --scope waits for the scope even with --no-block
(verified 2026-10-03), hence the backgrounding. Same fix in
chromebox-watchdog.sh (timers currently off).

dm.py: dm_read() now uses run_full() and prints a WARNING to stderr
with rc + last error line instead of failing silently on empty reads.

Trailers: Session: sidechat/opm-blind-fix
2026-10-04 01:16:59 +00:00
operator a7639e2341 propagation: single registry drives all node references
New bin/netvm-registry.py parses NODES.md (node -> cdp_port); every
consumer reads from it instead of hardcoding:
- onboard-driver.py: CDP_PORTS dict -> registry lookup (new nodes work)
- muse-signin.py: hardcoded 9410 -> --node/--cdp-port args
- muse-chat-api.py: hardcoded ACCOUNTS -> registry-built
- agent-health.sh: hardcoded muse/pip blocks -> loop over all active
  nodes (646 and opm now get health coverage too)
NODES.md: record opm node (9440).
2026-10-03 21:28:36 +00:00
operator ed45611912 Add GOLDEN PATH headers to agent loops 2026-10-03 18:04:57 +00:00
operator baae422b5f agent-health.sh: operator health monitor 2026-10-03 18:03:17 +00:00