Files
box/docs/CHROMEBOX-RUNBOOK.md
T
operator-main eba6c6965f Add CHROMEBOX-RUNBOOK.md: operator guide for the fleet browser stack
Covers architecture, watchdogs, queue, monitors, common failures, and 2am triage.

Session: sidechat/chromebox-ops
2026-10-04 13:27:29 +00:00

238 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Chromebox Runbook
Operator's guide to the Chromebox/NetVM browser fleet on `bl` (100.123.153.75).
Written 2026-10-04. If you're reading this at 2am, start at [Quick Triage](#quick-triage).
## Quick Triage
```bash
# 1. Are the browsers alive? (should return Browser JSON for all 4)
for t in 10.201.35.2:9410 10.201.87.2:9420 10.201.202.2:9430 10.201.157.2:9440; do
printf "%s: " $t; curl -s -m 5 -o /dev/null -w "%{http_code}\n" http://$t/json/version
done
# 2. Are the relay processes running? (expect 4 python3 netvm-cdp-relay.py)
pgrep -af netvm-cdp-relay | grep -v grep
# 3. What did the watchdogs do recently?
tail -20 /home/super/Projects/NetVM/chromebox-watchdog.log
tail -20 /home/super/Projects/NetVM/cdp-relay-watchdog.log
# 4. Are the host veth IPs assigned? (all 4 must show inet 10.201.x.1/30)
ip -o addr show | grep '10.201' | grep -v '10.201.202\|10.201.157' # muse+pip
ip -o addr show | grep -E '10.201.(35|87|157|202).1/30'
```
If step 1 shows 000 for a node but the browser is fine inside its netns,
it's almost certainly a **missing veth IP** or **dead relay** — see [Common Failures](#common-failures).
## Architecture
Four nodes, one per agent. 1:1 mapping: node == agent == profile == netns.
| Node | Agent | Netns | CDP Port | Veth IP (host) | Peer IP (netns) | Relay Target |
|------|-------|-------|----------|----------------|-----------------|--------------|
| muse | muse | warp-muse | 9410 | 10.201.35.1/30 | 10.201.35.2 | 10.201.35.2:9410 |
| pip | pip | warp-pip | 9420 | 10.201.87.1/30 | 10.201.87.2 | 10.201.87.2:9420 |
| 646 | operator-646 | warp-646 | 9430 | 10.201.202.1/30 | 10.201.202.2 | 10.201.202.2:9430 |
| opm | opm | warp-opm | 9440 | 10.201.157.1/30 | 10.201.157.2 | 10.201.157.2:9440 |
**Traffic path** (per node):
```
host → veth IP:port (e.g. 10.201.87.2:9420)
→ netvm-cdp-relay.py (runs INSIDE the netns, listens on veth/peer IP)
→ 127.0.0.1:port (Chromium's DevTools, bound to netns loopback only)
→ Chromium (headless, Warp egress via WireGuard in the same netns)
```
**Key facts that will save you hours:**
- The relay listens on the **veth/peer IP**, NOT on 127.0.0.1. Curling `127.0.0.1:9420`
on the host returns nothing — that's normal, not a failure.
- Veth IPs are hash-derived via `bin/netvm-names.sh` (`netvm_names <node>` gives
you VETH, GW, PEER_IP). Don't hardcode them in new scripts — source the file.
- CDP ports 9410–9440 are registry-pinned. The hash-derived `CDP_PORT` from
`netvm-names.sh` is WRONG unless `CDP_PORT_OVERRIDE` was set. Always use the
pinned ports above.
- Chromium binds DevTools to loopback only. The relay exists because iptables
REDIRECT to 127.0.0.1 was proven not to establish — the userspace relay is
the reliable path. Don't try to "simplify" it away.
## Watchdogs
### chromebox-watchdog (browser health)
- **Script:** `/home/super/Projects/NetVM/bin/chromebox-watchdog.sh`
- **Timers:** `chromebox-watchdog-<profile>.timer` (one per profile: muse, pip, 646, opm)
- **Cadence:** every 2 minutes
- **Log:** `/home/super/Projects/NetVM/chromebox-watchdog.log` (10 MB rotation, 1 backup gen)
- **Per-profile Chromium output:** `/home/super/Projects/NetVM/chromebox-<profile>.log`
**Health check stages** (in order, `HEALTH_FAIL_REASON` tells you which broke):
1. Chromium process exists for the profile → else `"no chromium process for profile"`
2. CDP responds on the profile's port → else `"CDP unreachable on :<port>"`
3. Chat page title present in target list → else reason names the missing title
**Kill-loop guard** (added 2026-10-04): before killing, checks if a Chromium for
the profile launched <2 min ago. If so, skips the kill — it's probably still
starting. Prevents the watchdog from murdering a slow-but-working browser.
**Relaunch retry** (added 2026-10-04): after relaunch, tries the health check up
to 4 times, 15s apart (~60s window). Only declares FAILED if all 4 miss.
A single 25s check was too brittle for cold starts (fresh egress IP +
Cloudflare handshake can exceed 25s).
**What "relaunch FAILED — needs operator attention" means:** the watchdog tried
4 times and the browser still isn't healthy. Check `HEALTH_FAIL_REASON` in the
log, then look at the per-profile chromium log. Don't just re-run the watchdog —
find out why the browser won't come up.
### cdp-relay-watchdog (relay health)
- **Script:** `/home/super/Projects/NetVM/bin/cdp-relay-watchdog.sh`
- **Timer:** `cdp-relay-watchdog.timer`
- **Service:** `cdp-relay-watchdog.service` (Type=oneshot)
- **Cadence:** every 5 minutes
- **Log:** `/home/super/Projects/NetVM/cdp-relay-watchdog.log`
**Two-stage check** (per node):
1. **Veth IP present** (`veth_healthy`): verifies the host veth interface has its
expected IP. If missing → `FAIL_LOUD` in the log, skips relay restart
(pointless without the veth). Does NOT auto-recreate the veth — that touches
WireGuard/iptables, too invasive for a watchdog.
2. **Relay connectivity** (`relay_healthy`): curls `http://<peer-ip>:<port>/json/version`
and greps for `"Browser"`. **Never trusts pidfiles** — they go stale and lie.
If a relay is down but the veth is fine, the watchdog kills any existing relay
for that port and restarts it with the exact `netvm-node-up.sh` invocation
(inside the netns, listening on the veth IP).
## The Queue (`cdp_queue.py`)
**Why it exists:** nothing coordinated browser operations. DM sends, tab opens,
agent reads, and watchdog restarts all hit the same Chromium with zero
scheduling. Under fleet concurrency this overwhelms bl's CPU.
**Location:** `/home/super/Projects/NetVM/bin/cdp_queue.py`
**Design:**
- Per-node FIFO queue (muse, pip, 646, opm each independent)
- Max **2 concurrent** CDP operations per browser (`MAX_CONCURRENT = 2`)
- Priority levels: `PRIORITY_HIGH = 0` (DM sends, user-facing),
`PRIORITY_NORMAL = 1` (default), `PRIORITY_LOW = 2`
- Cross-process coordination via **flock'd ticket files** in `/tmp/cdp-queue/<node>/`
(not just threading — `dm.py` shells out to `muse-chat-api.py`, so in-process
semaphores alone don't coordinate)
- `QueueTimeout` after 60s (configurable per-call) — fails loud, never hangs forever
- Warns when wait exceeds 10s
**Integration points:**
- `muse-chat-api.py`: every CDP session wrapped in `with cdp_slot(node, priority=...)`.
Priority from `CDP_PRIORITY` env var (`high`/`normal`/`low`, default `normal`).
Graceful fallback: if `cdp_queue` import fails, runs unqueued (nullcontext).
This is the single choke point — all current and future callers get queuing.
- `dm.py`: `run()` accepts `priority` param, sets `CDP_PRIORITY` in subprocess env.
`dm_send()` uses `priority="high"`. Read-back verification stays normal.
**Usage:**
```python
from cdp_queue import cdp_slot, PRIORITY_HIGH
with cdp_slot("opm", priority=PRIORITY_HIGH):
# ... CDP operations ...
```
**Monitoring:** `queue_depth(node)` returns current depth. `/tmp/cdp-queue/<node>/`
has `slot-0.lock` / `slot-1.lock` (persistent, by design) and transient ticket
files (should clean up within ~60s of completion — slow async cleanup, not a leak).
## Monitors (side-chat crons)
Five runtime crons report to the Chromebox ops side chat. All silent when healthy.
| Monitor | Cadence | What it checks | Alert threshold |
|---------|---------|----------------|-----------------|
| chromebox-watchdog-alerts | 10 min | New `relaunch FAILED` lines in watchdog log | Any new failure |
| cdp-relay-health-monitor | 15 min | All 4 relays return HTTP 200 on veth IPs | Any non-200 |
| browser-flap-detector | 30 min | `relaunch OK` count per profile in last hour | 3+ restarts/hour |
| cdp-latency-monitor | 15 min | CDP `/json/version` response time per node | FAIL or >5s (2–5s tracked, alerts after 3 consecutive) |
| chromebox-error-log-watch | 1 hour | FATAL/crash/segfault/OOM in per-profile chrome logs | Any new match |
Watermarks live in `~/workspace/goals/chromebox-ops/hidden_files/` so alerts
only fire on genuinely new events, not repeats.
## Common Failures and Fixes
### Missing host veth IP
**Symptoms:** relay process running, but host can't reach `http://<peer-ip>:<port>`.
`ip addr show dev ve-<tag>` shows no `inet 10.201.x.1/30`.
**Cause:** veth pair exists but host-side IP was lost (observed 2026-10-04 for
muse/pip — interfaces up, IPs gone).
**Fix:** `sudo ip addr add <GW>/30 dev <VETH>` where GW/VETH come from
`netvm-names.sh`. Or run `netvm-node-up.sh <node>` (idempotent, also fixes
anything else that's drifted).
**Detect:** cdp-relay-watchdog logs `FAIL_LOUD` for this. Don't just restart
the relay — it won't help without the veth IP.
### Stale pidfiles
**Symptoms:** `netvm-topology.sh` (old versions) reports relay DOWN when it's up,
or UP on the wrong port. `/run/netvm-<node>-cdp-relay.pid` points to a dead or
recycled PID.
**Fix:** fixed 2026-10-04 — topology script now checks by connectivity, not
pidfile. If you see pidfile-based checks anywhere else, they're suspect.
### Kill loop (watchdog killing slow browser)
**Symptoms:** repeated `unhealthy, relaunching` → `relaunch FAILED` cycles in
quick succession. Browser process exists and CDP responds, but the page hasn't
loaded yet.
**Cause (pre-2026-10-04):** single 25s post-launch check, no retry, no
recent-launch guard.
**Fix:** fixed 2026-10-04 — 4 retries over 60s + skip kill if launched <2 min ago.
If you see this pattern again, the guard may need tuning (longer window).
### Relay on wrong IP
**Symptoms:** relay process exists but on the wrong veth IP (e.g., muse's relay
on pip's `10.201.87.2` instead of muse's `10.201.35.2`). Port responds on the
wrong node's IP.
**Cause:** relay started with wrong PEER_IP argument.
**Fix:** kill it, restart with the correct IP from `netvm-names.sh`:
`sudo ip netns exec warp-<node> setsid nohup python3 bin/netvm-cdp-relay.py <PEER_IP> <PORT> 127.0.0.1 <PORT>`
### Queue module missing
**Symptoms:** everything works but nothing is actually queued — `/tmp/cdp-queue/`
never gets ticket files. No errors (graceful fallback hides it).
**Cause:** `bin/cdp_queue.py` not present on bl (observed 2026-10-04 — integration
was wired but the module file was missing).
**Fix:** ensure the file exists and `python3 -m py_compile` passes. It's committed
in the NetVM repo — check `git status` if it's gone.
### Orphan relays on random ports
**Symptoms:** `pgrep -af netvm-cdp-relay` shows relays on ports like 9269, 9278,
9353, 10239, 10355 (hash-derived, not registry ports).
**Cause:** old node-ups or queue tests. Harmless but confusing.
**Fix:** kill them. Only 9410/9420/9430/9440 should be running.
## Key Reference
**Node bring-up:** `sudo /home/super/Projects/NetVM/bin/netvm-node-up.sh <node>`
(idempotent — safe to re-run, recreates veth/WireGuard/relay as needed)
**Node teardown:** `bin/netvm-node-down.sh <node>`
**Topology check:** `bin/netvm-topology.sh` (connectivity-based since 2026-10-04)
**Names:** `bin/netvm-names.sh` — source it: `. bin/netvm-names.sh && netvm_names <node>`
gives you `$NETNS $WG $VETH $VPEER $SUB $GW $PEER_IP`
**DM send:** `python3 bin/dm.py send --agent <from> --to <to> --target main "message"`
**Commit identity** (shared repo — always use per-command flags, never repo config):
```bash
git -c user.name="operator-main" -c user.email="operator-main@frontdoor.local" commit ...
```
**What NOT to do:**
- Don't trust pidfiles for relay health — they lie.
- Don't curl `127.0.0.1:<port>` on the host to check relays — nothing listens there.
- Don't kill a browser that just launched — check `ps -o lstart` first.
- Don't commit other sessions' uncommitted work — use `git add -p` or selective staging.
- Don't restart browsers to fix relay problems — relays and browsers are independent.