Files
box/bin/agent-health.sh
T
operator-main f19aff9b79 Fix supervisor SIGKILL loop killing restarted browsers (opm/pip)
Root cause: agent-health.service runs Type=oneshot with the default
KillMode=control-group. restart_browser() spawned the replacement
chromium with nohup under the service, so systemd SIGKILLed it the
moment the service exited. Every 5-min tick: FAIL -> restart ->
RECOVERED -> SIGKILL at teardown. opm and pip were permanently dark
and dm.py read masked it (empty output, exit 0).

Fix: launch replacements via systemd-run --user --scope (backgrounded)
so the browser lives in a transient scope outside the service cgroup
and survives teardown. setsid does NOT escape either. Note:
systemd-run --scope waits for the scope even with --no-block
(verified 2026-10-03), hence the backgrounding. Same fix in
chromebox-watchdog.sh (timers currently off).

dm.py: dm_read() now uses run_full() and prints a WARNING to stderr
with rc + last error line instead of failing silently on empty reads.

Trailers: Session: sidechat/opm-blind-fix
2026-10-04 01:16:59 +00:00

86 lines
3.2 KiB
Bash
Executable File

#!/bin/bash
# GOLDEN PATH: container -> VM (34.139.37.135) -> bl (100.123.153.75) -> netns -> browser -> agent
# This script is the operator's heartbeat. If it stops, agents go dark.
# When debugging: trace each hop. Don't assume -- verify.
# Operator health monitor for Muse agents on bl.
# Checks each agent via API every 5 minutes. If unresponsive:
# 1. Restart the browser
# 2. Re-check
# 3. Log failure if still down
#
# Run via systemd timer or cron: */5 * * * * /home/super/Projects/NetVM/bin/agent-health.sh
#
# Agents are defined in ACCOUNTS.md. This script reads the active ones.
NETVM_BIN="$(cd "$(dirname "$0")" && pwd)"
LOG="/tmp/agent-health.log"
check_agent() {
local agent=$1
local node=$2
local cdp_port=$3
# Try API messages command (lightweight check)
if timeout 30 "$NETVM_BIN/netvm-exec.sh" "$node" -- python3 "$NETVM_BIN/muse-chat-api.py" --account "$agent" messages 1 > /dev/null 2>&1; then
echo "$(date -Iseconds) $agent: OK" >> "$LOG"
return 0
else
echo "$(date -Iseconds) $agent: FAIL (api timeout)" >> "$LOG"
return 1
fi
}
restart_browser() {
local agent=$1
local cdp_port=$2
echo "$(date -Iseconds) $agent: restarting browser..." >> "$LOG"
# Kill existing
pkill -f "$agent.*$cdp_port" 2>/dev/null
sleep 3
# Restart via netvm-chrome.sh in its own systemd scope.
# This oneshot service runs with KillMode=control-group, so anything
# spawned directly under it (nohup AND setsid both stay in the cgroup)
# is SIGKILLed when the service exits — observed 2026-10-03: every
# restart "recovered" then died at service teardown, looping forever.
# A transient scope escapes the service cgroup and survives.
# NOTE: systemd-run --scope WAITS for the scope's processes (even with
# --no-block, verified 2026-10-03), so background it — the scope itself
# is an independent unit and outlives the wrapper.
cd "$NETVM_BIN"
systemd-run --user --scope --unit="netvm-chrome-$agent-$(date +%s)" \
./netvm-chrome.sh --headless --cdp-port "$cdp_port" "$agent" https://muse.ai \
> "/tmp/bl-$agent.log" 2>&1 &
sleep 15
echo "$(date -Iseconds) $agent: browser restarted" >> "$LOG"
}
# Main
echo "=== Health check $(date -Iseconds) ===" >> "$LOG"
check_one() {
local agent=$1
local cdp_port=$2
# node == agent == profile (unified naming)
if ! check_agent "$agent" "$agent" "$cdp_port"; then
restart_browser "$agent" "$cdp_port"
sleep 5
if ! check_agent "$agent" "$agent" "$cdp_port"; then
echo "$(date -Iseconds) $agent: CRITICAL - still down after restart" >> "$LOG"
# TODO: Alert operator (e.g., via board post or email)
else
echo "$(date -Iseconds) $agent: RECOVERED after restart" >> "$LOG"
fi
fi
}
# Every active node in the NODES.md registry gets checked — new nodes
# propagate automatically, no per-node blocks to add.
"$NETVM_BIN/netvm-registry.py" 2>/dev/null | while IFS=: read -r node port; do
[ -n "$node" ] && [ -n "$port" ] && check_one "$node" "$port"
done
# Trim log (keep last 1000 lines)
tail -1000 "$LOG" > "$LOG.tmp" && mv "$LOG.tmp" "$LOG"