From eba6c6965f81949df00318820979014fecc32d51 Mon Sep 17 00:00:00 2001 From: operator-main Date: Sun, 4 Oct 2026 13:27:29 +0000 Subject: [PATCH] Add CHROMEBOX-RUNBOOK.md: operator guide for the fleet browser stack Covers architecture, watchdogs, queue, monitors, common failures, and 2am triage. Session: sidechat/chromebox-ops --- docs/CHROMEBOX-RUNBOOK.md | 237 ++++++++++++++++++++++++++++++++++++++ 1 file changed, 237 insertions(+) create mode 100644 docs/CHROMEBOX-RUNBOOK.md diff --git a/docs/CHROMEBOX-RUNBOOK.md b/docs/CHROMEBOX-RUNBOOK.md new file mode 100644 index 0000000..054a435 --- /dev/null +++ b/docs/CHROMEBOX-RUNBOOK.md @@ -0,0 +1,237 @@ +# Chromebox Runbook + +Operator's guide to the Chromebox/NetVM browser fleet on `bl` (100.123.153.75). +Written 2026-10-04. If you're reading this at 2am, start at [Quick Triage](#quick-triage). + +## Quick Triage + +```bash +# 1. Are the browsers alive? (should return Browser JSON for all 4) +for t in 10.201.35.2:9410 10.201.87.2:9420 10.201.202.2:9430 10.201.157.2:9440; do + printf "%s: " $t; curl -s -m 5 -o /dev/null -w "%{http_code}\n" http://$t/json/version +done + +# 2. Are the relay processes running? (expect 4 python3 netvm-cdp-relay.py) +pgrep -af netvm-cdp-relay | grep -v grep + +# 3. What did the watchdogs do recently? +tail -20 /home/super/Projects/NetVM/chromebox-watchdog.log +tail -20 /home/super/Projects/NetVM/cdp-relay-watchdog.log + +# 4. Are the host veth IPs assigned? (all 4 must show inet 10.201.x.1/30) +ip -o addr show | grep '10.201' | grep -v '10.201.202\|10.201.157' # muse+pip +ip -o addr show | grep -E '10.201.(35|87|157|202).1/30' +``` + +If step 1 shows 000 for a node but the browser is fine inside its netns, +it's almost certainly a **missing veth IP** or **dead relay** — see [Common Failures](#common-failures). + +## Architecture + +Four nodes, one per agent. 1:1 mapping: node == agent == profile == netns. + +| Node | Agent | Netns | CDP Port | Veth IP (host) | Peer IP (netns) | Relay Target | +|------|-------|-------|----------|----------------|-----------------|--------------| +| muse | muse | warp-muse | 9410 | 10.201.35.1/30 | 10.201.35.2 | 10.201.35.2:9410 | +| pip | pip | warp-pip | 9420 | 10.201.87.1/30 | 10.201.87.2 | 10.201.87.2:9420 | +| 646 | operator-646 | warp-646 | 9430 | 10.201.202.1/30 | 10.201.202.2 | 10.201.202.2:9430 | +| opm | opm | warp-opm | 9440 | 10.201.157.1/30 | 10.201.157.2 | 10.201.157.2:9440 | + +**Traffic path** (per node): +``` +host → veth IP:port (e.g. 10.201.87.2:9420) + → netvm-cdp-relay.py (runs INSIDE the netns, listens on veth/peer IP) + → 127.0.0.1:port (Chromium's DevTools, bound to netns loopback only) + → Chromium (headless, Warp egress via WireGuard in the same netns) +``` + +**Key facts that will save you hours:** +- The relay listens on the **veth/peer IP**, NOT on 127.0.0.1. Curling `127.0.0.1:9420` + on the host returns nothing — that's normal, not a failure. +- Veth IPs are hash-derived via `bin/netvm-names.sh` (`netvm_names ` gives + you VETH, GW, PEER_IP). Don't hardcode them in new scripts — source the file. +- CDP ports 9410–9440 are registry-pinned. The hash-derived `CDP_PORT` from + `netvm-names.sh` is WRONG unless `CDP_PORT_OVERRIDE` was set. Always use the + pinned ports above. +- Chromium binds DevTools to loopback only. The relay exists because iptables + REDIRECT to 127.0.0.1 was proven not to establish — the userspace relay is + the reliable path. Don't try to "simplify" it away. + +## Watchdogs + +### chromebox-watchdog (browser health) + +- **Script:** `/home/super/Projects/NetVM/bin/chromebox-watchdog.sh` +- **Timers:** `chromebox-watchdog-.timer` (one per profile: muse, pip, 646, opm) +- **Cadence:** every 2 minutes +- **Log:** `/home/super/Projects/NetVM/chromebox-watchdog.log` (10 MB rotation, 1 backup gen) +- **Per-profile Chromium output:** `/home/super/Projects/NetVM/chromebox-.log` + +**Health check stages** (in order, `HEALTH_FAIL_REASON` tells you which broke): +1. Chromium process exists for the profile → else `"no chromium process for profile"` +2. CDP responds on the profile's port → else `"CDP unreachable on :"` +3. Chat page title present in target list → else reason names the missing title + +**Kill-loop guard** (added 2026-10-04): before killing, checks if a Chromium for +the profile launched <2 min ago. If so, skips the kill — it's probably still +starting. Prevents the watchdog from murdering a slow-but-working browser. + +**Relaunch retry** (added 2026-10-04): after relaunch, tries the health check up +to 4 times, 15s apart (~60s window). Only declares FAILED if all 4 miss. +A single 25s check was too brittle for cold starts (fresh egress IP + +Cloudflare handshake can exceed 25s). + +**What "relaunch FAILED — needs operator attention" means:** the watchdog tried +4 times and the browser still isn't healthy. Check `HEALTH_FAIL_REASON` in the +log, then look at the per-profile chromium log. Don't just re-run the watchdog — +find out why the browser won't come up. + +### cdp-relay-watchdog (relay health) + +- **Script:** `/home/super/Projects/NetVM/bin/cdp-relay-watchdog.sh` +- **Timer:** `cdp-relay-watchdog.timer` +- **Service:** `cdp-relay-watchdog.service` (Type=oneshot) +- **Cadence:** every 5 minutes +- **Log:** `/home/super/Projects/NetVM/cdp-relay-watchdog.log` + +**Two-stage check** (per node): +1. **Veth IP present** (`veth_healthy`): verifies the host veth interface has its + expected IP. If missing → `FAIL_LOUD` in the log, skips relay restart + (pointless without the veth). Does NOT auto-recreate the veth — that touches + WireGuard/iptables, too invasive for a watchdog. +2. **Relay connectivity** (`relay_healthy`): curls `http://:/json/version` + and greps for `"Browser"`. **Never trusts pidfiles** — they go stale and lie. + +If a relay is down but the veth is fine, the watchdog kills any existing relay +for that port and restarts it with the exact `netvm-node-up.sh` invocation +(inside the netns, listening on the veth IP). + +## The Queue (`cdp_queue.py`) + +**Why it exists:** nothing coordinated browser operations. DM sends, tab opens, +agent reads, and watchdog restarts all hit the same Chromium with zero +scheduling. Under fleet concurrency this overwhelms bl's CPU. + +**Location:** `/home/super/Projects/NetVM/bin/cdp_queue.py` + +**Design:** +- Per-node FIFO queue (muse, pip, 646, opm each independent) +- Max **2 concurrent** CDP operations per browser (`MAX_CONCURRENT = 2`) +- Priority levels: `PRIORITY_HIGH = 0` (DM sends, user-facing), + `PRIORITY_NORMAL = 1` (default), `PRIORITY_LOW = 2` +- Cross-process coordination via **flock'd ticket files** in `/tmp/cdp-queue//` + (not just threading — `dm.py` shells out to `muse-chat-api.py`, so in-process + semaphores alone don't coordinate) +- `QueueTimeout` after 60s (configurable per-call) — fails loud, never hangs forever +- Warns when wait exceeds 10s + +**Integration points:** +- `muse-chat-api.py`: every CDP session wrapped in `with cdp_slot(node, priority=...)`. + Priority from `CDP_PRIORITY` env var (`high`/`normal`/`low`, default `normal`). + Graceful fallback: if `cdp_queue` import fails, runs unqueued (nullcontext). + This is the single choke point — all current and future callers get queuing. +- `dm.py`: `run()` accepts `priority` param, sets `CDP_PRIORITY` in subprocess env. + `dm_send()` uses `priority="high"`. Read-back verification stays normal. + +**Usage:** +```python +from cdp_queue import cdp_slot, PRIORITY_HIGH +with cdp_slot("opm", priority=PRIORITY_HIGH): + # ... CDP operations ... +``` + +**Monitoring:** `queue_depth(node)` returns current depth. `/tmp/cdp-queue//` +has `slot-0.lock` / `slot-1.lock` (persistent, by design) and transient ticket +files (should clean up within ~60s of completion — slow async cleanup, not a leak). + +## Monitors (side-chat crons) + +Five runtime crons report to the Chromebox ops side chat. All silent when healthy. + +| Monitor | Cadence | What it checks | Alert threshold | +|---------|---------|----------------|-----------------| +| chromebox-watchdog-alerts | 10 min | New `relaunch FAILED` lines in watchdog log | Any new failure | +| cdp-relay-health-monitor | 15 min | All 4 relays return HTTP 200 on veth IPs | Any non-200 | +| browser-flap-detector | 30 min | `relaunch OK` count per profile in last hour | 3+ restarts/hour | +| cdp-latency-monitor | 15 min | CDP `/json/version` response time per node | FAIL or >5s (2–5s tracked, alerts after 3 consecutive) | +| chromebox-error-log-watch | 1 hour | FATAL/crash/segfault/OOM in per-profile chrome logs | Any new match | + +Watermarks live in `~/workspace/goals/chromebox-ops/hidden_files/` so alerts +only fire on genuinely new events, not repeats. + +## Common Failures and Fixes + +### Missing host veth IP +**Symptoms:** relay process running, but host can't reach `http://:`. +`ip addr show dev ve-` shows no `inet 10.201.x.1/30`. +**Cause:** veth pair exists but host-side IP was lost (observed 2026-10-04 for +muse/pip — interfaces up, IPs gone). +**Fix:** `sudo ip addr add /30 dev ` where GW/VETH come from +`netvm-names.sh`. Or run `netvm-node-up.sh ` (idempotent, also fixes +anything else that's drifted). +**Detect:** cdp-relay-watchdog logs `FAIL_LOUD` for this. Don't just restart +the relay — it won't help without the veth IP. + +### Stale pidfiles +**Symptoms:** `netvm-topology.sh` (old versions) reports relay DOWN when it's up, +or UP on the wrong port. `/run/netvm--cdp-relay.pid` points to a dead or +recycled PID. +**Fix:** fixed 2026-10-04 — topology script now checks by connectivity, not +pidfile. If you see pidfile-based checks anywhere else, they're suspect. + +### Kill loop (watchdog killing slow browser) +**Symptoms:** repeated `unhealthy, relaunching` → `relaunch FAILED` cycles in +quick succession. Browser process exists and CDP responds, but the page hasn't +loaded yet. +**Cause (pre-2026-10-04):** single 25s post-launch check, no retry, no +recent-launch guard. +**Fix:** fixed 2026-10-04 — 4 retries over 60s + skip kill if launched <2 min ago. +If you see this pattern again, the guard may need tuning (longer window). + +### Relay on wrong IP +**Symptoms:** relay process exists but on the wrong veth IP (e.g., muse's relay +on pip's `10.201.87.2` instead of muse's `10.201.35.2`). Port responds on the +wrong node's IP. +**Cause:** relay started with wrong PEER_IP argument. +**Fix:** kill it, restart with the correct IP from `netvm-names.sh`: +`sudo ip netns exec warp- setsid nohup python3 bin/netvm-cdp-relay.py 127.0.0.1 ` + +### Queue module missing +**Symptoms:** everything works but nothing is actually queued — `/tmp/cdp-queue/` +never gets ticket files. No errors (graceful fallback hides it). +**Cause:** `bin/cdp_queue.py` not present on bl (observed 2026-10-04 — integration +was wired but the module file was missing). +**Fix:** ensure the file exists and `python3 -m py_compile` passes. It's committed +in the NetVM repo — check `git status` if it's gone. + +### Orphan relays on random ports +**Symptoms:** `pgrep -af netvm-cdp-relay` shows relays on ports like 9269, 9278, +9353, 10239, 10355 (hash-derived, not registry ports). +**Cause:** old node-ups or queue tests. Harmless but confusing. +**Fix:** kill them. Only 9410/9420/9430/9440 should be running. + +## Key Reference + +**Node bring-up:** `sudo /home/super/Projects/NetVM/bin/netvm-node-up.sh ` +(idempotent — safe to re-run, recreates veth/WireGuard/relay as needed) + +**Node teardown:** `bin/netvm-node-down.sh ` + +**Topology check:** `bin/netvm-topology.sh` (connectivity-based since 2026-10-04) + +**Names:** `bin/netvm-names.sh` — source it: `. bin/netvm-names.sh && netvm_names ` +gives you `$NETNS $WG $VETH $VPEER $SUB $GW $PEER_IP` + +**DM send:** `python3 bin/dm.py send --agent --to --target main "message"` + +**Commit identity** (shared repo — always use per-command flags, never repo config): +```bash +git -c user.name="operator-main" -c user.email="operator-main@frontdoor.local" commit ... +``` + +**What NOT to do:** +- Don't trust pidfiles for relay health — they lie. +- Don't curl `127.0.0.1:` on the host to check relays — nothing listens there. +- Don't kill a browser that just launched — check `ps -o lstart` first. +- Don't commit other sessions' uncommitted work — use `git add -p` or selective staging. +- Don't restart browsers to fix relay problems — relays and browsers are independent.