Files
box/docs/CHROMEBOX-RUNBOOK.md
T

281 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Chromebox Runbook
> **Box is the main surface.** All operator work goes through Box (box.muse-dev.online). The web UI, `box` CLI, and agents share the same API endpoints. No UI-only powers.
Operator's guide to the Chromebox/NetVM browser fleet on `bl` (100.123.153.75).
Written 2026-10-04. If you're reading this at 2am, start at [Quick Triage](#quick-triage).
## Quick Triage
```bash
# 1. Are the browsers alive? (should return Browser JSON for all 4)
for t in 10.201.35.2:9410 10.201.87.2:9420 10.201.202.2:9430 10.201.157.2:9440; do
printf "%s: " $t; curl -s -m 5 -o /dev/null -w "%{http_code}\n" http://$t/json/version
done
# 2. Are the relay processes running? (expect 4 python3 netvm-cdp-relay.py)
pgrep -af netvm-cdp-relay | grep -v grep
# 3. What did the watchdogs do recently?
tail -20 /home/super/Projects/NetVM/chromebox-watchdog.log
tail -20 /home/super/Projects/NetVM/cdp-relay-watchdog.log
# 4. Are the host veth IPs assigned? (all 4 must show inet 10.201.x.1/30)
ip -o addr show | grep '10.201' | grep -v '10.201.202\|10.201.157' # muse+pip
ip -o addr show | grep -E '10.201.(35|87|157|202).1/30'
```
If step 1 shows 000 for a node but the browser is fine inside its netns,
it's almost certainly a **missing veth IP** or **dead relay** — see [Common Failures](#common-failures).
## Architecture
Four nodes, one per agent. 1:1 mapping: node == agent == profile == netns.
| Node | Agent | Netns | CDP Port | Veth IP (host) | Peer IP (netns) | Relay Target |
|------|-------|-------|----------|----------------|-----------------|--------------|
| muse | muse | warp-muse | 9410 | 10.201.35.1/30 | 10.201.35.2 | 10.201.35.2:9410 |
| pip | pip | warp-pip | 9420 | 10.201.87.1/30 | 10.201.87.2 | 10.201.87.2:9420 |
| 646 | operator-646 | warp-646 | 9430 | 10.201.202.1/30 | 10.201.202.2 | 10.201.202.2:9430 |
| opm | opm | warp-opm | 9440 | 10.201.157.1/30 | 10.201.157.2 | 10.201.157.2:9440 |
**Traffic path** (per node):
```
host → veth IP:port (e.g. 10.201.87.2:9420)
→ netvm-cdp-relay.py (runs INSIDE the netns, listens on veth/peer IP)
→ 127.0.0.1:port (Chromium's DevTools, bound to netns loopback only)
→ Chromium (headless, Warp egress via WireGuard in the same netns)
```
**Key facts that will save you hours:**
- The relay listens on the **veth/peer IP**, NOT on 127.0.0.1. Curling `127.0.0.1:9420`
on the host returns nothing — that's normal, not a failure.
- Veth IPs are hash-derived via `bin/netvm-names.sh` (`netvm_names <node>` gives
you VETH, GW, PEER_IP). Don't hardcode them in new scripts — source the file.
- CDP ports 9410–9440 are registry-pinned. The hash-derived `CDP_PORT` from
`netvm-names.sh` is WRONG unless `CDP_PORT_OVERRIDE` was set. Always use the
pinned ports above.
- Chromium binds DevTools to loopback only. The relay exists because iptables
REDIRECT to 127.0.0.1 was proven not to establish — the userspace relay is
the reliable path. Don't try to "simplify" it away.
## Watchdogs
### chromebox-watchdog (browser health)
- **Script:** `/home/super/Projects/NetVM/bin/chromebox-watchdog.sh`
- **Timers:** `chromebox-watchdog-<profile>.timer` (one per profile — every active registry node: muse, pip, 646, opm, def, dev)
- **Cadence:** every 2 minutes
- **Log:** `/home/super/Projects/NetVM/chromebox-watchdog.log` (10 MB rotation, 1 backup gen)
- **Per-profile Chromium output:** `/home/super/Projects/NetVM/chromebox-<profile>.log`
**Health check stages** (in order, `HEALTH_FAIL_REASON` tells you which broke):
1. Chromium process exists for the profile → else `"no chromium process for profile"`
2. CDP responds on the profile's port → else `"CDP unreachable on :<port>"`
3. Chat page title present in target list → else reason names the missing title
**Kill-loop guard** (added 2026-10-04): before killing, checks if a Chromium for
the profile launched <2 min ago. If so, skips the kill — it's probably still
starting. Prevents the watchdog from murdering a slow-but-working browser.
**Relaunch retry** (added 2026-10-04): after relaunch, tries the health check up
to 4 times, 15s apart (~60s window). Only declares FAILED if all 4 miss.
A single 25s check was too brittle for cold starts (fresh egress IP +
Cloudflare handshake can exceed 25s).
**What "relaunch FAILED — needs operator attention" means:** the watchdog tried
4 times and the browser still isn't healthy. Check `HEALTH_FAIL_REASON` in the
log, then look at the per-profile chromium log. Don't just re-run the watchdog —
find out why the browser won't come up.
### cdp-relay-watchdog (relay health)
- **Script:** `/home/super/Projects/NetVM/bin/cdp-relay-watchdog.sh`
- **Timer:** `cdp-relay-watchdog.timer`
- **Service:** `cdp-relay-watchdog.service` (Type=oneshot)
- **Cadence:** every 5 minutes
- **Log:** `/home/super/Projects/NetVM/cdp-relay-watchdog.log`
**Two-stage check** (per node):
1. **Veth IP present** (`veth_healthy`): verifies the host veth interface has its
expected IP. If missing → `FAIL_LOUD` in the log, skips relay restart
(pointless without the veth). Does NOT auto-recreate the veth — that touches
WireGuard/iptables, too invasive for a watchdog.
2. **Relay connectivity** (`relay_healthy`): curls `http://<peer-ip>:<port>/json/version`
and greps for `"Browser"`. **Never trusts pidfiles** — they go stale and lie.
If a relay is down but the veth is fine, the watchdog kills any existing relay
for that port and restarts it with the exact `netvm-node-up.sh` invocation
(inside the netns, listening on the veth IP).
## The Queue (`cdp_queue.py`)
**Why it exists:** nothing coordinated browser operations. DM sends, tab opens,
agent reads, and watchdog restarts all hit the same Chromium with zero
scheduling. Under fleet concurrency this overwhelms bl's CPU.
**Location:** `/home/super/Projects/NetVM/bin/cdp_queue.py`
**Design:**
- Per-node FIFO queue (muse, pip, 646, opm each independent)
- Max **2 concurrent** CDP operations per browser (`MAX_CONCURRENT = 2`)
- Priority levels: `PRIORITY_HIGH = 0` (DM sends, user-facing),
`PRIORITY_NORMAL = 1` (default), `PRIORITY_LOW = 2`
- Cross-process coordination via **flock'd ticket files** in `/tmp/cdp-queue/<node>/`
(not just threading — `dm.py` shells out to `muse-chat-api.py`, so in-process
semaphores alone don't coordinate)
- `QueueTimeout` after 60s (configurable per-call) — fails loud, never hangs forever
- Warns when wait exceeds 10s
**Integration points:**
- `muse-chat-api.py`: every CDP session wrapped in `with cdp_slot(node, priority=...)`.
Priority from `CDP_PRIORITY` env var (`high`/`normal`/`low`, default `normal`).
Graceful fallback: if `cdp_queue` import fails, runs unqueued (nullcontext).
This is the single choke point — all current and future callers get queuing.
- `dm.py`: `run()` accepts `priority` param, sets `CDP_PRIORITY` in subprocess env.
`dm_send()` uses `priority="high"`. Read-back verification stays normal.
**Usage:**
```python
from cdp_queue import cdp_slot, PRIORITY_HIGH
with cdp_slot("opm", priority=PRIORITY_HIGH):
# ... CDP operations ...
```
**Monitoring:** `queue_depth(node)` returns current depth. `/tmp/cdp-queue/<node>/`
has `slot-0.lock` / `slot-1.lock` (persistent, by design) and transient ticket
files (should clean up within ~60s of completion — slow async cleanup, not a leak).
## Monitors (side-chat crons)
Five runtime crons report to the Chromebox ops side chat. All silent when healthy.
| Monitor | Cadence | What it checks | Alert threshold |
|---------|---------|----------------|-----------------|
| chromebox-watchdog-alerts | 10 min | New `relaunch FAILED` lines in watchdog log | Any new failure |
| cdp-relay-health-monitor | 15 min | All 4 relays return HTTP 200 on veth IPs | Any non-200 |
| browser-flap-detector | 30 min | `relaunch OK` count per profile in last hour | 3+ restarts/hour |
| cdp-latency-monitor | 15 min | CDP `/json/version` response time per node | FAIL or >5s (2–5s tracked, alerts after 3 consecutive) |
| chromebox-error-log-watch | 1 hour | FATAL/crash/segfault/OOM in per-profile chrome logs | Any new match |
Watermarks live in `~/workspace/goals/chromebox-ops/hidden_files/` so alerts
only fire on genuinely new events, not repeats.
## Common Failures and Fixes
### Missing host veth IP
**Symptoms:** relay process running, but host can't reach `http://<peer-ip>:<port>`.
`ip addr show dev ve-<tag>` shows no `inet 10.201.x.1/30`.
**Cause:** veth pair exists but host-side IP was lost (observed 2026-10-04 for
muse/pip — interfaces up, IPs gone).
**Fix:** `sudo ip addr add <GW>/30 dev <VETH>` where GW/VETH come from
`netvm-names.sh`. Or run `netvm-node-up.sh <node>` (idempotent, also fixes
anything else that's drifted).
**Detect:** cdp-relay-watchdog logs `FAIL_LOUD` for this. Don't just restart
the relay — it won't help without the veth IP.
### Stale pidfiles
**Symptoms:** `netvm-topology.sh` (old versions) reports relay DOWN when it's up,
or UP on the wrong port. `/run/netvm-<node>-cdp-relay.pid` points to a dead or
recycled PID.
**Fix:** fixed 2026-10-04 — topology script now checks by connectivity, not
pidfile. If you see pidfile-based checks anywhere else, they're suspect.
### Kill loop (watchdog killing slow browser)
**Symptoms:** repeated `unhealthy, relaunching` → `relaunch FAILED` cycles in
quick succession. Browser process exists and CDP responds, but the page hasn't
loaded yet.
**Cause (pre-2026-10-04):** single 25s post-launch check, no retry, no
recent-launch guard.
**Fix:** fixed 2026-10-04 — 4 retries over 60s + skip kill if launched <2 min ago.
If you see this pattern again, the guard may need tuning (longer window).
### Futile-restart loop (agent-health killing a healthy browser)
**Symptoms:** `FAIL (api timeout)` + `restarting browser...` + `CRITICAL -
still down after restart` repeating every ~10 min for one node while its
CDP port stays up (observed 2026-10-06: def, 57 restarts, 155 API FAILs).
The API/account layer is broken; restarts cannot fix it, they just murder
a working browser.
**Fix:** `agent-health.sh` circuit breaker — after 3 consecutive futile
restarts the circuit OPENS (alert in log + journal, no more kills) until
a 30-min half-open probe or any successful check. Manual reset:
`rm /tmp/agent-health-state/circuit-<node> /tmp/agent-health-state/futile-<node>`.
Then fix the actual API-layer failure (account session/auth/chat-state),
not the browser.
### Setup-fed supervision (new nodes automatically watched)
**Wiring:** `netvm-node-up.sh` ends with `ensure-node-supervision.sh <node>`
(idempotent): appends the NODES.md registry row (port from netvm-names
pinning, honors CDP_PORT_OVERRIDE) and installs/enables
`chromebox-watchdog-<node>.timer`. Provision/onboarding reach it
transitively via node-up. Registry-driven supervisors (relay watchdog,
agent-health, relay-health/cdp-latency checks) pick up new rows on
their next run — no per-node code edits. Heal drift anytime:
`sudo bin/ensure-node-supervision.sh --all`.
### Relay on wrong IP
**Symptoms:** relay process exists but on the wrong veth IP (e.g., muse's relay
on pip's `10.201.87.2` instead of muse's `10.201.35.2`). Port responds on the
wrong node's IP.
**Cause:** relay started with wrong PEER_IP argument.
**Fix:** kill it, restart with the correct IP from `netvm-names.sh`:
`sudo ip netns exec warp-<node> setsid nohup python3 bin/netvm-cdp-relay.py <PEER_IP> <PORT> 127.0.0.1 <PORT>`
### Queue module missing
**Symptoms:** everything works but nothing is actually queued — `/tmp/cdp-queue/`
never gets ticket files. No errors (graceful fallback hides it).
**Cause:** `bin/cdp_queue.py` not present on bl (observed 2026-10-04 — integration
was wired but the module file was missing).
**Fix:** ensure the file exists and `python3 -m py_compile` passes. It's committed
in the NetVM repo — check `git status` if it's gone.
### Orphan relays on random ports
**Symptoms:** `pgrep -af netvm-cdp-relay` shows relays on ports like 9269, 9278,
9353, 10239, 10355 (hash-derived, not registry ports).
**Cause:** old node-ups or queue tests. Harmless but confusing.
**Fix:** kill them. Only the registry ports (9410/9420/9430/9440/9450/9455) should be running.
## Fleet Status From Blind Shells
`box fleet status` probes live (pgrep + peer-IP CDP). Sandboxed shells
(own PID/net namespaces, no sudo, no route to 10.201.x.x) fail both
probes for every node. Instead of misreporting STOPPED, fleet status
falls back to host watchdog evidence (`bin/host_evidence.py`):
- Recent watchdog timer runs (journal) with no newer failure line in
`cdp-relay-watchdog.log` / `chromebox-watchdog.log` (both are
silent-when-healthy) prove the node is up → `ACTIVE [*]`.
- `UNKNOWN` means neither live probes nor host evidence could decide
(e.g. watchdog timers not installed yet for that node).
- Host evidence never overrides a live local signal, so a fresh outage
observed on the host always wins over a minutes-old watchdog run.
Same rule drives `box approvals check`: `BLIND` = CDP ok on host, this
shell cannot reach it; approval queues are unverified, not clear.
## Key Reference
**Node bring-up:** `sudo /home/super/Projects/NetVM/bin/netvm-node-up.sh <node>`
(idempotent — safe to re-run, recreates veth/WireGuard/relay as needed)
**Node teardown:** `bin/netvm-node-down.sh <node>`
**Topology check:** `bin/netvm-topology.sh` (connectivity-based since 2026-10-04)
**Names:** `bin/netvm-names.sh` — source it: `. bin/netvm-names.sh && netvm_names <node>`
gives you `$NETNS $WG $VETH $VPEER $SUB $GW $PEER_IP`
**DM send:** `python3 bin/dm.py send --agent <from> --to <to> --target main "message"`
**Commit identity** (shared repo — always use per-command flags, never repo config):
```bash
git -c user.name="operator-main" -c user.email="operator-main@frontdoor.local" commit ...
```
**What NOT to do:**
- Don't trust pidfiles for relay health — they lie.
- Don't curl `127.0.0.1:<port>` on the host to check relays — nothing listens there.
- Don't kill a browser that just launched — check `ps -o lstart` first.
- Don't commit other sessions' uncommitted work — use `git add -p` or selective staging.
- Don't restart browsers to fix relay problems — relays and browsers are independent.