Commit Graph

7 Commits

Author SHA1 Message Date
operator-main 76f854da6b Harden agent-health.sh: soften API-timeout kill path, add recent-relaunch guard
- Require 2 CONSECUTIVE muse-chat-api.py API failures before kill -9
  (per-node counter in /tmp/agent-health-state, reset on success).
  A single 30s API timeout killed 646s healthy browser at 20:48:44 UTC
  while its CDP port was still listening.
- Extend post-restart re-check grace to ~60s (15s internal + 45s), matching
  chromebox-watchdog.shs proven 60s retry window.
- Add recent_relaunch() guard (mirrors chromebox-watchdog.sh idiom):
  skip the kill path when the main browser process launched <2 min ago,
  so the two watchdogs can not kill each others fresh browsers.
- check_agent now returns 0/1/2 (healthy/api-fail/port-down); CDP-port
  failure still kills immediately. warp-$node checks untouched.

Session: sidechat/chromebox-fixes
2026-10-04 22:46:44 +00:00
operator-main 5ab157b3b0 fleet alerting: fix netns naming in agent-health.sh + add fleet-alert-check.sh
agent-health.sh used bare node names for 'ip netns exec' but netns are

named warp-<node> since the NetVM layout; the 6189793 CDP-liveness check

always failed ('No such file or directory'), logging false CRITICALs and

kill -9'ing healthy browsers every 5 min. Use warp-$node.

fleet-alert-check.sh: new 5-min critical-condition detector (per-node CDP

liveness via warp-<node> netns). Consecutive-failure state machine:

page after 2 consecutive failures, re-page every 30 min while critical,

quiet-hours-aware (first alert always pages). Emits ALERT/RECOVERY

records to ~/.local/share/fleet-alert/outbox.jsonl for the container

fleet-alert-relay hook; best-effort box-ctl notify to healthy agents.

Session: sidechat/critical-alerting-pipeline
2026-10-04 20:09:36 +00:00
operator-main 6189793748 agent-health.sh: Verify CDP port liveness, not just API success\n\nThe watchdog only checked if the API responded, missing zombie\nbrowsers (process alive but CDP not listening). Now explicitly\nchecks via ss -tln in the netns before trying the API.\nAlso: kill by exact PIDs instead of unreliable pkill patterns,\nand warn if port still bound after kill. 2026-10-04 02:23:20 +00:00
operator-main f19aff9b79 Fix supervisor SIGKILL loop killing restarted browsers (opm/pip)
Root cause: agent-health.service runs Type=oneshot with the default
KillMode=control-group. restart_browser() spawned the replacement
chromium with nohup under the service, so systemd SIGKILLed it the
moment the service exited. Every 5-min tick: FAIL -> restart ->
RECOVERED -> SIGKILL at teardown. opm and pip were permanently dark
and dm.py read masked it (empty output, exit 0).

Fix: launch replacements via systemd-run --user --scope (backgrounded)
so the browser lives in a transient scope outside the service cgroup
and survives teardown. setsid does NOT escape either. Note:
systemd-run --scope waits for the scope even with --no-block
(verified 2026-10-03), hence the backgrounding. Same fix in
chromebox-watchdog.sh (timers currently off).

dm.py: dm_read() now uses run_full() and prints a WARNING to stderr
with rc + last error line instead of failing silently on empty reads.

Trailers: Session: sidechat/opm-blind-fix
2026-10-04 01:16:59 +00:00
operator a7639e2341 propagation: single registry drives all node references
New bin/netvm-registry.py parses NODES.md (node -> cdp_port); every
consumer reads from it instead of hardcoding:
- onboard-driver.py: CDP_PORTS dict -> registry lookup (new nodes work)
- muse-signin.py: hardcoded 9410 -> --node/--cdp-port args
- muse-chat-api.py: hardcoded ACCOUNTS -> registry-built
- agent-health.sh: hardcoded muse/pip blocks -> loop over all active
  nodes (646 and opm now get health coverage too)
NODES.md: record opm node (9440).
2026-10-03 21:28:36 +00:00
operator ed45611912 Add GOLDEN PATH headers to agent loops 2026-10-03 18:04:57 +00:00
operator baae422b5f agent-health.sh: operator health monitor 2026-10-03 18:03:17 +00:00