3 Commits

Author SHA1 Message Date
operator 34b0ef9fe2 chore(fleet): sync operator memory, hatch menu dialogs, and watchdog alerts 2026-10-07 00:25:51 +00:00
operator-main ae1bf17de5 Fix relay watchdog false-positive: sudo the pkill
restart_relay() ran pkill without sudo, but relay processes are
root-owned and the timer runs as User=super. The kill failed EPERM
(silently swallowed by || true), the old relay kept running, and
SO_REUSEADDR let the replacement double-bind the same port. The
post-restart health check then passed and logged "restarted OK"
when nothing was actually restarted.

Adding sudo -n to the pkill, matching the sudo -n ip netns exec
already used to start the relay.

Session: sidechat/chromebox-ops
2026-10-04 18:22:10 +00:00
operator-main 343920e0c3 Watchdog upgrades: stage-specific logging + new CDP relay watchdog
chromebox-watchdog.sh: HEALTH_FAIL_REASON pinpoints which health stage failed (no process / CDP unreachable / no Chat page); chromium stdout redirected to per-profile chromebox-<profile>.log; 10MB log rotation (one generation); Chat title match relaxed to .*Chat.

bin/cdp-relay-watchdog.sh (new): keeps per-node CDP relays alive. Two-stage check: (1) host veth IP assigned (fail-loud, no auto-fix — veth recreation touches WireGuard/iptables), (2) relay connectivity via curl to veth IP:port (never trust pidfiles — observed stale 2026-10-04). Restarts dead/misrouted relays in-netns. Runs via systemd timer every 5min. Pattern mirrors chromebox-watchdog.sh.

Session: sidechat/chromebox-ops
2026-10-04 12:40:56 +00:00