10 Commits

Author SHA1 Message Date
operator 0c6d2235ab feat(kpi): add autonomous worker auto-spawn engine and watchdog reconciliation 2026-10-06 23:03:52 +00:00
operator-main 345eb09559 fix(watchdog): eliminate SIGPIPE+pipefail phantom failures in health checks
echo "$list" | grep -q under set -o pipefail exits 141 whenever grep
matches before echo finishes writing, so healthy browsers were reported
'CDP up but no page target' and killed every 2 min fleet-wide (load 19+).
Use [[ == *glob* ]] (no pipe, no race) for the page/muse.ai stages, and
grep -c (reads to EOF, never early-exits) for the netns check.
2026-10-06 07:28:20 +00:00
operator 4f4b8768b6 feat(alert): wire idempotent fleet-alert relay into check loop and stop stale chrome scopes 2026-10-05 17:41:04 +00:00
operator 1a271b1bbd feat(hybrid-gateway): integrate muse-cli with Cloudflare netns isolation, symmetric sidechat routing, and 646-pip sync unblock 2026-10-04 22:54:01 +00:00
operator-main c07802dfa2 Fix watchdog relaunch-loop: guard before relaunch, main-process PID filter
The relaunch-loop guard only protected the kill step, not the relaunch.
When CDP was unreachable on a slow-starting browser, the watchdog would
invoke netvm-chrome.sh (which kills the existing browser) before checking
if it was recently launched — piling up 5 chromiums on opm.

Now the <120s check runs before the relaunch and skips the entire cycle.
Also fixed the PID check to match only the main browser process
(--remote-debugging-port, excluding --type= renderer/gpu children).

Session: sidechat/chromebox-ops
2026-10-04 18:21:55 +00:00
operator-main cf8f59b737 fix(watchdog): generalize CDP page health check for Muse targets 2026-10-04 16:39:30 +00:00
operator 7a35b684c4 feat: unified fleet CLI, Main Chat preservation policy, sidechat routing, and file transfers
- Added CHAT_POLICY.md and README.md banner enforcing sidechat-first and file-transfer-first rules.
- Added strict Main Chat block to super dm send and super dm wo with --allow-main-chat override.
- Implemented file transfer staging and metadata registry in super dm send-file and super dm files (with clean subcommand).
- Added full job lifecycle management (show, create, enable, disable, delete, run --follow) to super-cli.py and box-ctl.py.
- Audited all jobs in jobs/*.json and redirected automated dispatches away from Main Chat.
- Hardened chromebox-watchdog.sh with systemd user session environment exports and stale singleton cleanup.
- Added compose_check command choice to muse-chat-api.py.
2026-10-04 16:34:25 +00:00
operator-main a32670731f chromebox-watchdog: fix kill loop on slow cold starts
Two fixes: (1) retry health check 4x with 15s gaps after relaunch instead of single 25s check; (2) skip kill if browser launched <2min ago (probably still starting). Prevents watchdog from killing a working-but-slow browser.

Session: sidechat/chromebox-ops
2026-10-04 13:17:13 +00:00
operator-main 343920e0c3 Watchdog upgrades: stage-specific logging + new CDP relay watchdog
chromebox-watchdog.sh: HEALTH_FAIL_REASON pinpoints which health stage failed (no process / CDP unreachable / no Chat page); chromium stdout redirected to per-profile chromebox-<profile>.log; 10MB log rotation (one generation); Chat title match relaxed to .*Chat.

bin/cdp-relay-watchdog.sh (new): keeps per-node CDP relays alive. Two-stage check: (1) host veth IP assigned (fail-loud, no auto-fix — veth recreation touches WireGuard/iptables), (2) relay connectivity via curl to veth IP:port (never trust pidfiles — observed stale 2026-10-04). Restarts dead/misrouted relays in-netns. Runs via systemd timer every 5min. Pattern mirrors chromebox-watchdog.sh.

Session: sidechat/chromebox-ops
2026-10-04 12:40:56 +00:00
operator-main f19aff9b79 Fix supervisor SIGKILL loop killing restarted browsers (opm/pip)
Root cause: agent-health.service runs Type=oneshot with the default
KillMode=control-group. restart_browser() spawned the replacement
chromium with nohup under the service, so systemd SIGKILLed it the
moment the service exited. Every 5-min tick: FAIL -> restart ->
RECOVERED -> SIGKILL at teardown. opm and pip were permanently dark
and dm.py read masked it (empty output, exit 0).

Fix: launch replacements via systemd-run --user --scope (backgrounded)
so the browser lives in a transient scope outside the service cgroup
and survives teardown. setsid does NOT escape either. Note:
systemd-run --scope waits for the scope even with --no-block
(verified 2026-10-03), hence the backgrounding. Same fix in
chromebox-watchdog.sh (timers currently off).

dm.py: dm_read() now uses run_full() and prints a WARNING to stderr
with rc + last error line instead of failing silently on empty reads.

Trailers: Session: sidechat/opm-blind-fix
2026-10-04 01:16:59 +00:00