2026-10-04 13:27:29 +00:00
# Chromebox Runbook
2026-10-05 15:58:37 +00:00
> **Box is the main surface.** All operator work goes through Box (box.muse-dev.online). The web UI, `box` CLI, and agents share the same API endpoints. No UI-only powers.
2026-10-04 13:27:29 +00:00
Operator's guide to the Chromebox/NetVM browser fleet on `bl` (100.123.153.75).
Written 2026-10-04. If you're reading this at 2am, start at [Quick Triage ](#quick-triage ).
## Quick Triage
``` bash
# 1. Are the browsers alive? (should return Browser JSON for all 4)
for t in 10.201.35.2:9410 10.201.87.2:9420 10.201.202.2:9430 10.201.157.2:9440; do
printf "%s: " $t ; curl -s -m 5 -o /dev/null -w "%{http_code}\n" http://$t /json/version
done
# 2. Are the relay processes running? (expect 4 python3 netvm-cdp-relay.py)
pgrep -af netvm-cdp-relay | grep -v grep
# 3. What did the watchdogs do recently?
tail -20 /home/super/Projects/NetVM/chromebox-watchdog.log
tail -20 /home/super/Projects/NetVM/cdp-relay-watchdog.log
# 4. Are the host veth IPs assigned? (all 4 must show inet 10.201.x.1/30)
ip -o addr show | grep '10.201' | grep -v '10.201.202\|10.201.157' # muse+pip
ip -o addr show | grep -E '10.201.(35|87|157|202).1/30'
```
If step 1 shows 000 for a node but the browser is fine inside its netns,
it's almost certainly a **missing veth IP ** or **dead relay ** — see [Common Failures ](#common-failures ).
## Architecture
Four nodes, one per agent. 1:1 mapping: node == agent == profile == netns.
| Node | Agent | Netns | CDP Port | Veth IP (host) | Peer IP (netns) | Relay Target |
|------|-------|-------|----------|----------------|-----------------|--------------|
| muse | muse | warp-muse | 9410 | 10.201.35.1/30 | 10.201.35.2 | 10.201.35.2:9410 |
| pip | pip | warp-pip | 9420 | 10.201.87.1/30 | 10.201.87.2 | 10.201.87.2:9420 |
| 646 | operator-646 | warp-646 | 9430 | 10.201.202.1/30 | 10.201.202.2 | 10.201.202.2:9430 |
| opm | opm | warp-opm | 9440 | 10.201.157.1/30 | 10.201.157.2 | 10.201.157.2:9440 |
**Traffic path ** (per node):
```
host → veth IP:port (e.g. 10.201.87.2:9420)
→ netvm-cdp-relay.py (runs INSIDE the netns, listens on veth/peer IP)
→ 127.0.0.1:port (Chromium's DevTools, bound to netns loopback only)
→ Chromium (headless, Warp egress via WireGuard in the same netns)
```
**Key facts that will save you hours: **
- The relay listens on the **veth/peer IP ** , NOT on 127.0.0.1. Curling `127.0.0.1:9420`
on the host returns nothing — that's normal, not a failure.
- Veth IPs are hash-derived via `bin/netvm-names.sh` (`netvm_names <node>` gives
you VETH, GW, PEER_IP). Don't hardcode them in new scripts — source the file.
- CDP ports 9410– 9440 are registry-pinned. The hash-derived `CDP_PORT` from
`netvm-names.sh` is WRONG unless `CDP_PORT_OVERRIDE` was set. Always use the
pinned ports above.
- Chromium binds DevTools to loopback only. The relay exists because iptables
REDIRECT to 127.0.0.1 was proven not to establish — the userspace relay is
the reliable path. Don't try to "simplify" it away.
## Watchdogs
### chromebox-watchdog (browser health)
- **Script:** `/home/super/Projects/NetVM/bin/chromebox-watchdog.sh`
2026-10-07 00:25:51 +00:00
- **Timers:** `chromebox-watchdog-<profile>.timer` (one per profile — every active registry node: muse, pip, 646, opm, def, dev)
2026-10-04 13:27:29 +00:00
- **Cadence:** every 2 minutes
- **Log:** `/home/super/Projects/NetVM/chromebox-watchdog.log` (10 MB rotation, 1 backup gen)
- **Per-profile Chromium output:** `/home/super/Projects/NetVM/chromebox-<profile>.log`
**Health check stages ** (in order, `HEALTH_FAIL_REASON` tells you which broke):
1. Chromium process exists for the profile → else `"no chromium process for profile"`
2. CDP responds on the profile's port → else `"CDP unreachable on :<port>"`
3. Chat page title present in target list → else reason names the missing title
**Kill-loop guard ** (added 2026-10-04): before killing, checks if a Chromium for
the profile launched <2 min ago. If so, skips the kill — it's probably still
starting. Prevents the watchdog from murdering a slow-but-working browser.
**Relaunch retry ** (added 2026-10-04): after relaunch, tries the health check up
to 4 times, 15s apart (~60s window). Only declares FAILED if all 4 miss.
A single 25s check was too brittle for cold starts (fresh egress IP +
Cloudflare handshake can exceed 25s).
**What "relaunch FAILED — needs operator attention" means: ** the watchdog tried
4 times and the browser still isn't healthy. Check `HEALTH_FAIL_REASON` in the
log, then look at the per-profile chromium log. Don't just re-run the watchdog —
find out why the browser won't come up.
### cdp-relay-watchdog (relay health)
- **Script:** `/home/super/Projects/NetVM/bin/cdp-relay-watchdog.sh`
- **Timer:** `cdp-relay-watchdog.timer`
- **Service:** `cdp-relay-watchdog.service` (Type=oneshot)
- **Cadence:** every 5 minutes
- **Log:** `/home/super/Projects/NetVM/cdp-relay-watchdog.log`
**Two-stage check ** (per node):
1. **Veth IP present ** (`veth_healthy` ): verifies the host veth interface has its
expected IP. If missing → `FAIL_LOUD` in the log, skips relay restart
(pointless without the veth). Does NOT auto-recreate the veth — that touches
WireGuard/iptables, too invasive for a watchdog.
2. **Relay connectivity ** (`relay_healthy` ): curls `http://<peer-ip>:<port>/json/version`
and greps for `"Browser"` . **Never trusts pidfiles ** — they go stale and lie.
If a relay is down but the veth is fine, the watchdog kills any existing relay
for that port and restarts it with the exact `netvm-node-up.sh` invocation
(inside the netns, listening on the veth IP).
## The Queue (`cdp_queue.py`)
**Why it exists: ** nothing coordinated browser operations. DM sends, tab opens,
agent reads, and watchdog restarts all hit the same Chromium with zero
scheduling. Under fleet concurrency this overwhelms bl's CPU.
**Location: ** `/home/super/Projects/NetVM/bin/cdp_queue.py`
**Design: **
- Per-node FIFO queue (muse, pip, 646, opm each independent)
- Max **2 concurrent ** CDP operations per browser (`MAX_CONCURRENT = 2` )
- Priority levels: `PRIORITY_HIGH = 0` (DM sends, user-facing),
`PRIORITY_NORMAL = 1` (default), `PRIORITY_LOW = 2`
- Cross-process coordination via **flock'd ticket files ** in `/tmp/cdp-queue/<node>/`
(not just threading — `dm.py` shells out to `muse-chat-api.py` , so in-process
semaphores alone don't coordinate)
- `QueueTimeout` after 60s (configurable per-call) — fails loud, never hangs forever
- Warns when wait exceeds 10s
**Integration points: **
- `muse-chat-api.py` : every CDP session wrapped in `with cdp_slot(node, priority=...)` .
Priority from `CDP_PRIORITY` env var (`high` /`normal` /`low` , default `normal` ).
Graceful fallback: if `cdp_queue` import fails, runs unqueued (nullcontext).
This is the single choke point — all current and future callers get queuing.
- `dm.py` : `run()` accepts `priority` param, sets `CDP_PRIORITY` in subprocess env.
`dm_send()` uses `priority="high"` . Read-back verification stays normal.
**Usage: **
``` python
from cdp_queue import cdp_slot , PRIORITY_HIGH
with cdp_slot ( " opm " , priority = PRIORITY_HIGH ) :
# ... CDP operations ...
```
**Monitoring: ** `queue_depth(node)` returns current depth. `/tmp/cdp-queue/<node>/`
has `slot-0.lock` / `slot-1.lock` (persistent, by design) and transient ticket
files (should clean up within ~60s of completion — slow async cleanup, not a leak).
## Monitors (side-chat crons)
Five runtime crons report to the Chromebox ops side chat. All silent when healthy.
| Monitor | Cadence | What it checks | Alert threshold |
|---------|---------|----------------|-----------------|
| chromebox-watchdog-alerts | 10 min | New `relaunch FAILED` lines in watchdog log | Any new failure |
| cdp-relay-health-monitor | 15 min | All 4 relays return HTTP 200 on veth IPs | Any non-200 |
| browser-flap-detector | 30 min | `relaunch OK` count per profile in last hour | 3+ restarts/hour |
| cdp-latency-monitor | 15 min | CDP `/json/version` response time per node | FAIL or >5s (2– 5s tracked, alerts after 3 consecutive) |
| chromebox-error-log-watch | 1 hour | FATAL/crash/segfault/OOM in per-profile chrome logs | Any new match |
Watermarks live in `~/workspace/goals/chromebox-ops/hidden_files/` so alerts
only fire on genuinely new events, not repeats.
## Common Failures and Fixes
### Missing host veth IP
**Symptoms:** relay process running, but host can't reach `http://<peer-ip>:<port>` .
`ip addr show dev ve-<tag>` shows no `inet 10.201.x.1/30` .
**Cause: ** veth pair exists but host-side IP was lost (observed 2026-10-04 for
muse/pip — interfaces up, IPs gone).
**Fix: ** `sudo ip addr add <GW>/30 dev <VETH>` where GW/VETH come from
`netvm-names.sh` . Or run `netvm-node-up.sh <node>` (idempotent, also fixes
anything else that's drifted).
**Detect: ** cdp-relay-watchdog logs `FAIL_LOUD` for this. Don't just restart
the relay — it won't help without the veth IP.
### Stale pidfiles
**Symptoms:** `netvm-topology.sh` (old versions) reports relay DOWN when it's up,
or UP on the wrong port. `/run/netvm-<node>-cdp-relay.pid` points to a dead or
recycled PID.
**Fix: ** fixed 2026-10-04 — topology script now checks by connectivity, not
pidfile. If you see pidfile-based checks anywhere else, they're suspect.
### Kill loop (watchdog killing slow browser)
**Symptoms:** repeated `unhealthy, relaunching` → `relaunch FAILED` cycles in
quick succession. Browser process exists and CDP responds, but the page hasn't
loaded yet.
**Cause (pre-2026-10-04): ** single 25s post-launch check, no retry, no
recent-launch guard.
**Fix: ** fixed 2026-10-04 — 4 retries over 60s + skip kill if launched <2 min ago.
If you see this pattern again, the guard may need tuning (longer window).
2026-10-06 18:16:14 +00:00
### Futile-restart loop (agent-health killing a healthy browser)
**Symptoms:** `FAIL (api timeout)` + `restarting browser...` + `CRITICAL -
still down after restart` repeating every ~10 min for one node while its
CDP port stays up (observed 2026-10-06: def, 57 restarts, 155 API FAILs).
The API/account layer is broken; restarts cannot fix it, they just murder
a working browser.
**Fix: ** `agent-health.sh` circuit breaker — after 3 consecutive futile
restarts the circuit OPENS (alert in log + journal, no more kills) until
a 30-min half-open probe or any successful check. Manual reset:
`rm /tmp/agent-health-state/circuit-<node> /tmp/agent-health-state/futile-<node>` .
Then fix the actual API-layer failure (account session/auth/chat-state),
not the browser.
2026-10-06 19:28:29 +00:00
### Setup-fed supervision (new nodes automatically watched)
**Wiring:** `netvm-node-up.sh` ends with `ensure-node-supervision.sh <node>`
(idempotent): appends the NODES.md registry row (port from netvm-names
pinning, honors CDP_PORT_OVERRIDE) and installs/enables
`chromebox-watchdog-<node>.timer` . Provision/onboarding reach it
transitively via node-up. Registry-driven supervisors (relay watchdog,
agent-health, relay-health/cdp-latency checks) pick up new rows on
their next run — no per-node code edits. Heal drift anytime:
`sudo bin/ensure-node-supervision.sh --all` .
2026-10-04 13:27:29 +00:00
### Relay on wrong IP
**Symptoms:** relay process exists but on the wrong veth IP (e.g., muse's relay
on pip's `10.201.87.2` instead of muse's `10.201.35.2` ). Port responds on the
wrong node's IP.
**Cause: ** relay started with wrong PEER_IP argument.
**Fix: ** kill it, restart with the correct IP from `netvm-names.sh` :
`sudo ip netns exec warp-<node> setsid nohup python3 bin/netvm-cdp-relay.py <PEER_IP> <PORT> 127.0.0.1 <PORT>`
### Queue module missing
**Symptoms:** everything works but nothing is actually queued — `/tmp/cdp-queue/`
never gets ticket files. No errors (graceful fallback hides it).
**Cause: ** `bin/cdp_queue.py` not present on bl (observed 2026-10-04 — integration
was wired but the module file was missing).
**Fix: ** ensure the file exists and `python3 -m py_compile` passes. It's committed
in the NetVM repo — check `git status` if it's gone.
### Orphan relays on random ports
**Symptoms:** `pgrep -af netvm-cdp-relay` shows relays on ports like 9269, 9278,
9353, 10239, 10355 (hash-derived, not registry ports).
**Cause: ** old node-ups or queue tests. Harmless but confusing.
2026-10-07 00:25:51 +00:00
**Fix: ** kill them. Only the registry ports (9410/9420/9430/9440/9450/9455) should be running.
2026-10-04 13:27:29 +00:00
2026-10-06 18:16:14 +00:00
## Fleet Status From Blind Shells
`box fleet status` probes live (pgrep + peer-IP CDP). Sandboxed shells
(own PID/net namespaces, no sudo, no route to 10.201.x.x) fail both
probes for every node. Instead of misreporting STOPPED, fleet status
falls back to host watchdog evidence (`bin/host_evidence.py` ):
- Recent watchdog timer runs (journal) with no newer failure line in
`cdp-relay-watchdog.log` / `chromebox-watchdog.log` (both are
silent-when-healthy) prove the node is up → `ACTIVE [*]` .
- `UNKNOWN` means neither live probes nor host evidence could decide
2026-10-07 00:25:51 +00:00
(e.g. watchdog timers not installed yet for that node).
2026-10-06 18:16:14 +00:00
- Host evidence never overrides a live local signal, so a fresh outage
observed on the host always wins over a minutes-old watchdog run.
Same rule drives `box approvals check` : `BLIND` = CDP ok on host, this
shell cannot reach it; approval queues are unverified, not clear.
2026-10-04 13:27:29 +00:00
## Key Reference
**Node bring-up: ** `sudo /home/super/Projects/NetVM/bin/netvm-node-up.sh <node>`
(idempotent — safe to re-run, recreates veth/WireGuard/relay as needed)
**Node teardown: ** `bin/netvm-node-down.sh <node>`
**Topology check: ** `bin/netvm-topology.sh` (connectivity-based since 2026-10-04)
**Names: ** `bin/netvm-names.sh` — source it: `. bin/netvm-names.sh && netvm_names <node>`
gives you `$NETNS $WG $VETH $VPEER $SUB $GW $PEER_IP`
**DM send: ** `python3 bin/dm.py send --agent <from> --to <to> --target main "message"`
**Commit identity ** (shared repo — always use per-command flags, never repo config):
``` bash
git -c user.name= "operator-main" -c user.email= "operator-main@frontdoor.local" commit ...
```
**What NOT to do: **
- Don't trust pidfiles for relay health — they lie.
- Don't curl `127.0.0.1:<port>` on the host to check relays — nothing listens there.
- Don't kill a browser that just launched — check `ps -o lstart` first.
- Don't commit other sessions' uncommitted work — use `git add -p` or selective staging.
- Don't restart browsers to fix relay problems — relays and browsers are independent.