Files
box/docs/CHROMEBOX-RUNBOOK.md
T
Muse Sidechat 66c8900a58 feat: setup-fed watchdog supervision for all registry nodes
Close the def/dev supervision gap at the source: every node brought
up gets watched, and every supervisor enumerates the registry.

- bin/ensure-node-supervision.sh (new, idempotent): appends the
  NODES.md row (netvm-names port, honors CDP_PORT_OVERRIDE so it
  never fights provision's picker) and installs/enables
  chromebox-watchdog-<node>.timer. --all heals drift (registry +
  /etc/netvm identities). Template verified byte-identical to the
  installed def unit.
- netvm-node-up.sh: calls ensure (non-fatal) at the end. Provision
  and the onboarding pipeline reach it transitively.
- relay-health-check.sh, cdp-latency-check.sh: registry-driven
  watched_nodes() + LIB_ONLY guards (were hardcoded 4 nodes).
- tests/test_node_supervision.py (6): row add/idempotent/override,
  timer render, node-up wiring, both watched_nodes().
- CHROMEBOX-RUNBOOK.md: setup-fed supervision section.

Pairs with the registry-driven relay/chromebox watchdogs: new rows
are picked up on the next run with no per-node code edits.
2026-10-06 19:28:29 +00:00

14 KiB
Raw Blame History

Chromebox Runbook

Box is the main surface. All operator work goes through Box (box.muse-dev.online). The web UI, box CLI, and agents share the same API endpoints. No UI-only powers.

Operator's guide to the Chromebox/NetVM browser fleet on bl (100.123.153.75). Written 2026-10-04. If you're reading this at 2am, start at Quick Triage.

Quick Triage

# 1. Are the browsers alive? (should return Browser JSON for all 4)
for t in 10.201.35.2:9410 10.201.87.2:9420 10.201.202.2:9430 10.201.157.2:9440; do
  printf "%s: " $t; curl -s -m 5 -o /dev/null -w "%{http_code}\n" http://$t/json/version
done

# 2. Are the relay processes running? (expect 4 python3 netvm-cdp-relay.py)
pgrep -af netvm-cdp-relay | grep -v grep

# 3. What did the watchdogs do recently?
tail -20 /home/super/Projects/NetVM/chromebox-watchdog.log
tail -20 /home/super/Projects/NetVM/cdp-relay-watchdog.log

# 4. Are the host veth IPs assigned? (all 4 must show inet 10.201.x.1/30)
ip -o addr show | grep '10.201' | grep -v '10.201.202\|10.201.157'  # muse+pip
ip -o addr show | grep -E '10.201.(35|87|157|202).1/30'

If step 1 shows 000 for a node but the browser is fine inside its netns, it's almost certainly a missing veth IP or dead relay — see Common Failures.

Architecture

Four nodes, one per agent. 1:1 mapping: node == agent == profile == netns.

Node Agent Netns CDP Port Veth IP (host) Peer IP (netns) Relay Target
muse muse warp-muse 9410 10.201.35.1/30 10.201.35.2 10.201.35.2:9410
pip pip warp-pip 9420 10.201.87.1/30 10.201.87.2 10.201.87.2:9420
646 operator-646 warp-646 9430 10.201.202.1/30 10.201.202.2 10.201.202.2:9430
opm opm warp-opm 9440 10.201.157.1/30 10.201.157.2 10.201.157.2:9440

Traffic path (per node):

host → veth IP:port (e.g. 10.201.87.2:9420)
  → netvm-cdp-relay.py (runs INSIDE the netns, listens on veth/peer IP)
  → 127.0.0.1:port (Chromium's DevTools, bound to netns loopback only)
  → Chromium (headless, Warp egress via WireGuard in the same netns)

Key facts that will save you hours:

  • The relay listens on the veth/peer IP, NOT on 127.0.0.1. Curling 127.0.0.1:9420 on the host returns nothing — that's normal, not a failure.
  • Veth IPs are hash-derived via bin/netvm-names.sh (netvm_names <node> gives you VETH, GW, PEER_IP). Don't hardcode them in new scripts — source the file.
  • CDP ports 9410–9440 are registry-pinned. The hash-derived CDP_PORT from netvm-names.sh is WRONG unless CDP_PORT_OVERRIDE was set. Always use the pinned ports above.
  • Chromium binds DevTools to loopback only. The relay exists because iptables REDIRECT to 127.0.0.1 was proven not to establish — the userspace relay is the reliable path. Don't try to "simplify" it away.

Watchdogs

chromebox-watchdog (browser health)

  • Script: /home/super/Projects/NetVM/bin/chromebox-watchdog.sh
  • Timers: chromebox-watchdog-<profile>.timer (one per profile: muse, pip, 646, opm)
  • Cadence: every 2 minutes
  • Log: /home/super/Projects/NetVM/chromebox-watchdog.log (10 MB rotation, 1 backup gen)
  • Per-profile Chromium output: /home/super/Projects/NetVM/chromebox-<profile>.log

Health check stages (in order, HEALTH_FAIL_REASON tells you which broke):

  1. Chromium process exists for the profile → else "no chromium process for profile"
  2. CDP responds on the profile's port → else "CDP unreachable on :<port>"
  3. Chat page title present in target list → else reason names the missing title

Kill-loop guard (added 2026-10-04): before killing, checks if a Chromium for the profile launched <2 min ago. If so, skips the kill — it's probably still starting. Prevents the watchdog from murdering a slow-but-working browser.

Relaunch retry (added 2026-10-04): after relaunch, tries the health check up to 4 times, 15s apart (~60s window). Only declares FAILED if all 4 miss. A single 25s check was too brittle for cold starts (fresh egress IP + Cloudflare handshake can exceed 25s).

What "relaunch FAILED — needs operator attention" means: the watchdog tried 4 times and the browser still isn't healthy. Check HEALTH_FAIL_REASON in the log, then look at the per-profile chromium log. Don't just re-run the watchdog — find out why the browser won't come up.

cdp-relay-watchdog (relay health)

  • Script: /home/super/Projects/NetVM/bin/cdp-relay-watchdog.sh
  • Timer: cdp-relay-watchdog.timer
  • Service: cdp-relay-watchdog.service (Type=oneshot)
  • Cadence: every 5 minutes
  • Log: /home/super/Projects/NetVM/cdp-relay-watchdog.log

Two-stage check (per node):

  1. Veth IP present (veth_healthy): verifies the host veth interface has its expected IP. If missing → FAIL_LOUD in the log, skips relay restart (pointless without the veth). Does NOT auto-recreate the veth — that touches WireGuard/iptables, too invasive for a watchdog.
  2. Relay connectivity (relay_healthy): curls http://<peer-ip>:<port>/json/version and greps for "Browser". Never trusts pidfiles — they go stale and lie.

If a relay is down but the veth is fine, the watchdog kills any existing relay for that port and restarts it with the exact netvm-node-up.sh invocation (inside the netns, listening on the veth IP).

The Queue (cdp_queue.py)

Why it exists: nothing coordinated browser operations. DM sends, tab opens, agent reads, and watchdog restarts all hit the same Chromium with zero scheduling. Under fleet concurrency this overwhelms bl's CPU.

Location: /home/super/Projects/NetVM/bin/cdp_queue.py

Design:

  • Per-node FIFO queue (muse, pip, 646, opm each independent)
  • Max 2 concurrent CDP operations per browser (MAX_CONCURRENT = 2)
  • Priority levels: PRIORITY_HIGH = 0 (DM sends, user-facing), PRIORITY_NORMAL = 1 (default), PRIORITY_LOW = 2
  • Cross-process coordination via flock'd ticket files in /tmp/cdp-queue/<node>/ (not just threading — dm.py shells out to muse-chat-api.py, so in-process semaphores alone don't coordinate)
  • QueueTimeout after 60s (configurable per-call) — fails loud, never hangs forever
  • Warns when wait exceeds 10s

Integration points:

  • muse-chat-api.py: every CDP session wrapped in with cdp_slot(node, priority=...). Priority from CDP_PRIORITY env var (high/normal/low, default normal). Graceful fallback: if cdp_queue import fails, runs unqueued (nullcontext). This is the single choke point — all current and future callers get queuing.
  • dm.py: run() accepts priority param, sets CDP_PRIORITY in subprocess env. dm_send() uses priority="high". Read-back verification stays normal.

Usage:

from cdp_queue import cdp_slot, PRIORITY_HIGH
with cdp_slot("opm", priority=PRIORITY_HIGH):
    # ... CDP operations ...

Monitoring: queue_depth(node) returns current depth. /tmp/cdp-queue/<node>/ has slot-0.lock / slot-1.lock (persistent, by design) and transient ticket files (should clean up within ~60s of completion — slow async cleanup, not a leak).

Monitors (side-chat crons)

Five runtime crons report to the Chromebox ops side chat. All silent when healthy.

Monitor Cadence What it checks Alert threshold
chromebox-watchdog-alerts 10 min New relaunch FAILED lines in watchdog log Any new failure
cdp-relay-health-monitor 15 min All 4 relays return HTTP 200 on veth IPs Any non-200
browser-flap-detector 30 min relaunch OK count per profile in last hour 3+ restarts/hour
cdp-latency-monitor 15 min CDP /json/version response time per node FAIL or >5s (2–5s tracked, alerts after 3 consecutive)
chromebox-error-log-watch 1 hour FATAL/crash/segfault/OOM in per-profile chrome logs Any new match

Watermarks live in ~/workspace/goals/chromebox-ops/hidden_files/ so alerts only fire on genuinely new events, not repeats.

Common Failures and Fixes

Missing host veth IP

Symptoms: relay process running, but host can't reach http://<peer-ip>:<port>. ip addr show dev ve-<tag> shows no inet 10.201.x.1/30. Cause: veth pair exists but host-side IP was lost (observed 2026-10-04 for muse/pip — interfaces up, IPs gone). Fix: sudo ip addr add <GW>/30 dev <VETH> where GW/VETH come from netvm-names.sh. Or run netvm-node-up.sh <node> (idempotent, also fixes anything else that's drifted). Detect: cdp-relay-watchdog logs FAIL_LOUD for this. Don't just restart the relay — it won't help without the veth IP.

Stale pidfiles

Symptoms: netvm-topology.sh (old versions) reports relay DOWN when it's up, or UP on the wrong port. /run/netvm-<node>-cdp-relay.pid points to a dead or recycled PID. Fix: fixed 2026-10-04 — topology script now checks by connectivity, not pidfile. If you see pidfile-based checks anywhere else, they're suspect.

Kill loop (watchdog killing slow browser)

Symptoms: repeated unhealthy, relaunching → relaunch FAILED cycles in quick succession. Browser process exists and CDP responds, but the page hasn't loaded yet. Cause (pre-2026-10-04): single 25s post-launch check, no retry, no recent-launch guard. Fix: fixed 2026-10-04 — 4 retries over 60s + skip kill if launched <2 min ago. If you see this pattern again, the guard may need tuning (longer window).

Futile-restart loop (agent-health killing a healthy browser)

Symptoms: FAIL (api timeout) + restarting browser... + CRITICAL - still down after restart repeating every ~10 min for one node while its CDP port stays up (observed 2026-10-06: def, 57 restarts, 155 API FAILs). The API/account layer is broken; restarts cannot fix it, they just murder a working browser. Fix: agent-health.sh circuit breaker — after 3 consecutive futile restarts the circuit OPENS (alert in log + journal, no more kills) until a 30-min half-open probe or any successful check. Manual reset: rm /tmp/agent-health-state/circuit-<node> /tmp/agent-health-state/futile-<node>. Then fix the actual API-layer failure (account session/auth/chat-state), not the browser.

Setup-fed supervision (new nodes automatically watched)

Wiring: netvm-node-up.sh ends with ensure-node-supervision.sh <node> (idempotent): appends the NODES.md registry row (port from netvm-names pinning, honors CDP_PORT_OVERRIDE) and installs/enables chromebox-watchdog-<node>.timer. Provision/onboarding reach it transitively via node-up. Registry-driven supervisors (relay watchdog, agent-health, relay-health/cdp-latency checks) pick up new rows on their next run — no per-node code edits. Heal drift anytime: sudo bin/ensure-node-supervision.sh --all.

Relay on wrong IP

Symptoms: relay process exists but on the wrong veth IP (e.g., muse's relay on pip's 10.201.87.2 instead of muse's 10.201.35.2). Port responds on the wrong node's IP. Cause: relay started with wrong PEER_IP argument. Fix: kill it, restart with the correct IP from netvm-names.sh: sudo ip netns exec warp-<node> setsid nohup python3 bin/netvm-cdp-relay.py <PEER_IP> <PORT> 127.0.0.1 <PORT>

Queue module missing

Symptoms: everything works but nothing is actually queued — /tmp/cdp-queue/ never gets ticket files. No errors (graceful fallback hides it). Cause: bin/cdp_queue.py not present on bl (observed 2026-10-04 — integration was wired but the module file was missing). Fix: ensure the file exists and python3 -m py_compile passes. It's committed in the NetVM repo — check git status if it's gone.

Orphan relays on random ports

Symptoms: pgrep -af netvm-cdp-relay shows relays on ports like 9269, 9278, 9353, 10239, 10355 (hash-derived, not registry ports). Cause: old node-ups or queue tests. Harmless but confusing. Fix: kill them. Only 9410/9420/9430/9440 should be running.

Fleet Status From Blind Shells

box fleet status probes live (pgrep + peer-IP CDP). Sandboxed shells (own PID/net namespaces, no sudo, no route to 10.201.x.x) fail both probes for every node. Instead of misreporting STOPPED, fleet status falls back to host watchdog evidence (bin/host_evidence.py):

  • Recent watchdog timer runs (journal) with no newer failure line in cdp-relay-watchdog.log / chromebox-watchdog.log (both are silent-when-healthy) prove the node is up → ACTIVE [*].
  • UNKNOWN means neither live probes nor host evidence could decide (e.g. def/dev have no relay-monitor coverage).
  • Host evidence never overrides a live local signal, so a fresh outage observed on the host always wins over a minutes-old watchdog run.

Same rule drives box approvals check: BLIND = CDP ok on host, this shell cannot reach it; approval queues are unverified, not clear.

Key Reference

Node bring-up: sudo /home/super/Projects/NetVM/bin/netvm-node-up.sh <node> (idempotent — safe to re-run, recreates veth/WireGuard/relay as needed)

Node teardown: bin/netvm-node-down.sh <node>

Topology check: bin/netvm-topology.sh (connectivity-based since 2026-10-04)

Names: bin/netvm-names.sh — source it: . bin/netvm-names.sh && netvm_names <node> gives you $NETNS $WG $VETH $VPEER $SUB $GW $PEER_IP

DM send: python3 bin/dm.py send --agent <from> --to <to> --target main "message"

Commit identity (shared repo — always use per-command flags, never repo config):

git -c user.name="operator-main" -c user.email="operator-main@frontdoor.local" commit ...

What NOT to do:

  • Don't trust pidfiles for relay health — they lie.
  • Don't curl 127.0.0.1:<port> on the host to check relays — nothing listens there.
  • Don't kill a browser that just launched — check ps -o lstart first.
  • Don't commit other sessions' uncommitted work — use git add -p or selective staging.
  • Don't restart browsers to fix relay problems — relays and browsers are independent.