Files
box/docs/CHROMEBOX-RUNBOOK.md
T
Muse Sidechat c9143a558b fix: truthful fleet status in blind shells + agent-health circuit breaker
box fleet status / approvals check misreported every node as STOPPED /
CDP-unreachable from sandboxed shells (own PID+net namespaces: pgrep
blind, no route to 10.201.x.x, no sudo). Fleet was healthy throughout.

- bin/host_evidence.py (new): host watchdog evidence fallback. Recent
  timer runs (journal -o json, exact UNIT match) with no newer failure
  line in cdp-relay-watchdog.log / chromebox-watchdog.log (both
  silent-when-healthy) prove a node is up. def/dev have no watchdog
  coverage: browser verdict via chromebox-<node>.log freshness
  (alive-only), CDP verdict unknown.
- super-cli.py: effective status/source/evidence per node. Host
  evidence decides ONLY the fully-blind pattern (both local probes
  negative); live local signals always win. New UNKNOWN badge, [*]
  footnote; approvals UNREACHABLE splits into BLIND / OFFLINE(host
  agrees) / unreachable-evidence-inconclusive, with honest footer.
  proc_alive/cdp_ok keep local-probe meaning; status/source/evidence
  are new JSON fields.
- approvals.py: host_cdp_ok flag on the unreachable path.
- agent-health.sh: restart circuit breaker. 3 consecutive futile
  restarts (restart leaves agent still failing) opens the circuit:
  no more kills for 1800s, ALERT to log+journal, half-open probe
  after cooldown, reset on any success. Stops the def murder loop
  (57 restarts / 155 API FAILs for an account-layer failure).
- tests/test_fleet_status.py (25), tests/test_agent_health.py (6).
- CHROMEBOX-RUNBOOK.md: blind-shell status + futile-restart sections.

Tests: 98/98 focused green (agent_health + fleet_status +
completion + tool_calls). Live-verified: 4 ACTIVE [*] + 2 UNKNOWN.
2026-10-06 18:16:14 +00:00

271 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Chromebox Runbook
> **Box is the main surface.** All operator work goes through Box (box.muse-dev.online). The web UI, `box` CLI, and agents share the same API endpoints. No UI-only powers.
Operator's guide to the Chromebox/NetVM browser fleet on `bl` (100.123.153.75).
Written 2026-10-04. If you're reading this at 2am, start at [Quick Triage](#quick-triage).
## Quick Triage
```bash
# 1. Are the browsers alive? (should return Browser JSON for all 4)
for t in 10.201.35.2:9410 10.201.87.2:9420 10.201.202.2:9430 10.201.157.2:9440; do
printf "%s: " $t; curl -s -m 5 -o /dev/null -w "%{http_code}\n" http://$t/json/version
done
# 2. Are the relay processes running? (expect 4 python3 netvm-cdp-relay.py)
pgrep -af netvm-cdp-relay | grep -v grep
# 3. What did the watchdogs do recently?
tail -20 /home/super/Projects/NetVM/chromebox-watchdog.log
tail -20 /home/super/Projects/NetVM/cdp-relay-watchdog.log
# 4. Are the host veth IPs assigned? (all 4 must show inet 10.201.x.1/30)
ip -o addr show | grep '10.201' | grep -v '10.201.202\|10.201.157' # muse+pip
ip -o addr show | grep -E '10.201.(35|87|157|202).1/30'
```
If step 1 shows 000 for a node but the browser is fine inside its netns,
it's almost certainly a **missing veth IP** or **dead relay** — see [Common Failures](#common-failures).
## Architecture
Four nodes, one per agent. 1:1 mapping: node == agent == profile == netns.
| Node | Agent | Netns | CDP Port | Veth IP (host) | Peer IP (netns) | Relay Target |
|------|-------|-------|----------|----------------|-----------------|--------------|
| muse | muse | warp-muse | 9410 | 10.201.35.1/30 | 10.201.35.2 | 10.201.35.2:9410 |
| pip | pip | warp-pip | 9420 | 10.201.87.1/30 | 10.201.87.2 | 10.201.87.2:9420 |
| 646 | operator-646 | warp-646 | 9430 | 10.201.202.1/30 | 10.201.202.2 | 10.201.202.2:9430 |
| opm | opm | warp-opm | 9440 | 10.201.157.1/30 | 10.201.157.2 | 10.201.157.2:9440 |
**Traffic path** (per node):
```
host → veth IP:port (e.g. 10.201.87.2:9420)
→ netvm-cdp-relay.py (runs INSIDE the netns, listens on veth/peer IP)
→ 127.0.0.1:port (Chromium's DevTools, bound to netns loopback only)
→ Chromium (headless, Warp egress via WireGuard in the same netns)
```
**Key facts that will save you hours:**
- The relay listens on the **veth/peer IP**, NOT on 127.0.0.1. Curling `127.0.0.1:9420`
on the host returns nothing — that's normal, not a failure.
- Veth IPs are hash-derived via `bin/netvm-names.sh` (`netvm_names <node>` gives
you VETH, GW, PEER_IP). Don't hardcode them in new scripts — source the file.
- CDP ports 9410–9440 are registry-pinned. The hash-derived `CDP_PORT` from
`netvm-names.sh` is WRONG unless `CDP_PORT_OVERRIDE` was set. Always use the
pinned ports above.
- Chromium binds DevTools to loopback only. The relay exists because iptables
REDIRECT to 127.0.0.1 was proven not to establish — the userspace relay is
the reliable path. Don't try to "simplify" it away.
## Watchdogs
### chromebox-watchdog (browser health)
- **Script:** `/home/super/Projects/NetVM/bin/chromebox-watchdog.sh`
- **Timers:** `chromebox-watchdog-<profile>.timer` (one per profile: muse, pip, 646, opm)
- **Cadence:** every 2 minutes
- **Log:** `/home/super/Projects/NetVM/chromebox-watchdog.log` (10 MB rotation, 1 backup gen)
- **Per-profile Chromium output:** `/home/super/Projects/NetVM/chromebox-<profile>.log`
**Health check stages** (in order, `HEALTH_FAIL_REASON` tells you which broke):
1. Chromium process exists for the profile → else `"no chromium process for profile"`
2. CDP responds on the profile's port → else `"CDP unreachable on :<port>"`
3. Chat page title present in target list → else reason names the missing title
**Kill-loop guard** (added 2026-10-04): before killing, checks if a Chromium for
the profile launched <2 min ago. If so, skips the kill — it's probably still
starting. Prevents the watchdog from murdering a slow-but-working browser.
**Relaunch retry** (added 2026-10-04): after relaunch, tries the health check up
to 4 times, 15s apart (~60s window). Only declares FAILED if all 4 miss.
A single 25s check was too brittle for cold starts (fresh egress IP +
Cloudflare handshake can exceed 25s).
**What "relaunch FAILED — needs operator attention" means:** the watchdog tried
4 times and the browser still isn't healthy. Check `HEALTH_FAIL_REASON` in the
log, then look at the per-profile chromium log. Don't just re-run the watchdog —
find out why the browser won't come up.
### cdp-relay-watchdog (relay health)
- **Script:** `/home/super/Projects/NetVM/bin/cdp-relay-watchdog.sh`
- **Timer:** `cdp-relay-watchdog.timer`
- **Service:** `cdp-relay-watchdog.service` (Type=oneshot)
- **Cadence:** every 5 minutes
- **Log:** `/home/super/Projects/NetVM/cdp-relay-watchdog.log`
**Two-stage check** (per node):
1. **Veth IP present** (`veth_healthy`): verifies the host veth interface has its
expected IP. If missing → `FAIL_LOUD` in the log, skips relay restart
(pointless without the veth). Does NOT auto-recreate the veth — that touches
WireGuard/iptables, too invasive for a watchdog.
2. **Relay connectivity** (`relay_healthy`): curls `http://<peer-ip>:<port>/json/version`
and greps for `"Browser"`. **Never trusts pidfiles** — they go stale and lie.
If a relay is down but the veth is fine, the watchdog kills any existing relay
for that port and restarts it with the exact `netvm-node-up.sh` invocation
(inside the netns, listening on the veth IP).
## The Queue (`cdp_queue.py`)
**Why it exists:** nothing coordinated browser operations. DM sends, tab opens,
agent reads, and watchdog restarts all hit the same Chromium with zero
scheduling. Under fleet concurrency this overwhelms bl's CPU.
**Location:** `/home/super/Projects/NetVM/bin/cdp_queue.py`
**Design:**
- Per-node FIFO queue (muse, pip, 646, opm each independent)
- Max **2 concurrent** CDP operations per browser (`MAX_CONCURRENT = 2`)
- Priority levels: `PRIORITY_HIGH = 0` (DM sends, user-facing),
`PRIORITY_NORMAL = 1` (default), `PRIORITY_LOW = 2`
- Cross-process coordination via **flock'd ticket files** in `/tmp/cdp-queue/<node>/`
(not just threading — `dm.py` shells out to `muse-chat-api.py`, so in-process
semaphores alone don't coordinate)
- `QueueTimeout` after 60s (configurable per-call) — fails loud, never hangs forever
- Warns when wait exceeds 10s
**Integration points:**
- `muse-chat-api.py`: every CDP session wrapped in `with cdp_slot(node, priority=...)`.
Priority from `CDP_PRIORITY` env var (`high`/`normal`/`low`, default `normal`).
Graceful fallback: if `cdp_queue` import fails, runs unqueued (nullcontext).
This is the single choke point — all current and future callers get queuing.
- `dm.py`: `run()` accepts `priority` param, sets `CDP_PRIORITY` in subprocess env.
`dm_send()` uses `priority="high"`. Read-back verification stays normal.
**Usage:**
```python
from cdp_queue import cdp_slot, PRIORITY_HIGH
with cdp_slot("opm", priority=PRIORITY_HIGH):
# ... CDP operations ...
```
**Monitoring:** `queue_depth(node)` returns current depth. `/tmp/cdp-queue/<node>/`
has `slot-0.lock` / `slot-1.lock` (persistent, by design) and transient ticket
files (should clean up within ~60s of completion — slow async cleanup, not a leak).
## Monitors (side-chat crons)
Five runtime crons report to the Chromebox ops side chat. All silent when healthy.
| Monitor | Cadence | What it checks | Alert threshold |
|---------|---------|----------------|-----------------|
| chromebox-watchdog-alerts | 10 min | New `relaunch FAILED` lines in watchdog log | Any new failure |
| cdp-relay-health-monitor | 15 min | All 4 relays return HTTP 200 on veth IPs | Any non-200 |
| browser-flap-detector | 30 min | `relaunch OK` count per profile in last hour | 3+ restarts/hour |
| cdp-latency-monitor | 15 min | CDP `/json/version` response time per node | FAIL or >5s (2–5s tracked, alerts after 3 consecutive) |
| chromebox-error-log-watch | 1 hour | FATAL/crash/segfault/OOM in per-profile chrome logs | Any new match |
Watermarks live in `~/workspace/goals/chromebox-ops/hidden_files/` so alerts
only fire on genuinely new events, not repeats.
## Common Failures and Fixes
### Missing host veth IP
**Symptoms:** relay process running, but host can't reach `http://<peer-ip>:<port>`.
`ip addr show dev ve-<tag>` shows no `inet 10.201.x.1/30`.
**Cause:** veth pair exists but host-side IP was lost (observed 2026-10-04 for
muse/pip — interfaces up, IPs gone).
**Fix:** `sudo ip addr add <GW>/30 dev <VETH>` where GW/VETH come from
`netvm-names.sh`. Or run `netvm-node-up.sh <node>` (idempotent, also fixes
anything else that's drifted).
**Detect:** cdp-relay-watchdog logs `FAIL_LOUD` for this. Don't just restart
the relay — it won't help without the veth IP.
### Stale pidfiles
**Symptoms:** `netvm-topology.sh` (old versions) reports relay DOWN when it's up,
or UP on the wrong port. `/run/netvm-<node>-cdp-relay.pid` points to a dead or
recycled PID.
**Fix:** fixed 2026-10-04 — topology script now checks by connectivity, not
pidfile. If you see pidfile-based checks anywhere else, they're suspect.
### Kill loop (watchdog killing slow browser)
**Symptoms:** repeated `unhealthy, relaunching` → `relaunch FAILED` cycles in
quick succession. Browser process exists and CDP responds, but the page hasn't
loaded yet.
**Cause (pre-2026-10-04):** single 25s post-launch check, no retry, no
recent-launch guard.
**Fix:** fixed 2026-10-04 — 4 retries over 60s + skip kill if launched <2 min ago.
If you see this pattern again, the guard may need tuning (longer window).
### Futile-restart loop (agent-health killing a healthy browser)
**Symptoms:** `FAIL (api timeout)` + `restarting browser...` + `CRITICAL -
still down after restart` repeating every ~10 min for one node while its
CDP port stays up (observed 2026-10-06: def, 57 restarts, 155 API FAILs).
The API/account layer is broken; restarts cannot fix it, they just murder
a working browser.
**Fix:** `agent-health.sh` circuit breaker — after 3 consecutive futile
restarts the circuit OPENS (alert in log + journal, no more kills) until
a 30-min half-open probe or any successful check. Manual reset:
`rm /tmp/agent-health-state/circuit-<node> /tmp/agent-health-state/futile-<node>`.
Then fix the actual API-layer failure (account session/auth/chat-state),
not the browser.
### Relay on wrong IP
**Symptoms:** relay process exists but on the wrong veth IP (e.g., muse's relay
on pip's `10.201.87.2` instead of muse's `10.201.35.2`). Port responds on the
wrong node's IP.
**Cause:** relay started with wrong PEER_IP argument.
**Fix:** kill it, restart with the correct IP from `netvm-names.sh`:
`sudo ip netns exec warp-<node> setsid nohup python3 bin/netvm-cdp-relay.py <PEER_IP> <PORT> 127.0.0.1 <PORT>`
### Queue module missing
**Symptoms:** everything works but nothing is actually queued — `/tmp/cdp-queue/`
never gets ticket files. No errors (graceful fallback hides it).
**Cause:** `bin/cdp_queue.py` not present on bl (observed 2026-10-04 — integration
was wired but the module file was missing).
**Fix:** ensure the file exists and `python3 -m py_compile` passes. It's committed
in the NetVM repo — check `git status` if it's gone.
### Orphan relays on random ports
**Symptoms:** `pgrep -af netvm-cdp-relay` shows relays on ports like 9269, 9278,
9353, 10239, 10355 (hash-derived, not registry ports).
**Cause:** old node-ups or queue tests. Harmless but confusing.
**Fix:** kill them. Only 9410/9420/9430/9440 should be running.
## Fleet Status From Blind Shells
`box fleet status` probes live (pgrep + peer-IP CDP). Sandboxed shells
(own PID/net namespaces, no sudo, no route to 10.201.x.x) fail both
probes for every node. Instead of misreporting STOPPED, fleet status
falls back to host watchdog evidence (`bin/host_evidence.py`):
- Recent watchdog timer runs (journal) with no newer failure line in
`cdp-relay-watchdog.log` / `chromebox-watchdog.log` (both are
silent-when-healthy) prove the node is up → `ACTIVE [*]`.
- `UNKNOWN` means neither live probes nor host evidence could decide
(e.g. def/dev have no relay-monitor coverage).
- Host evidence never overrides a live local signal, so a fresh outage
observed on the host always wins over a minutes-old watchdog run.
Same rule drives `box approvals check`: `BLIND` = CDP ok on host, this
shell cannot reach it; approval queues are unverified, not clear.
## Key Reference
**Node bring-up:** `sudo /home/super/Projects/NetVM/bin/netvm-node-up.sh <node>`
(idempotent — safe to re-run, recreates veth/WireGuard/relay as needed)
**Node teardown:** `bin/netvm-node-down.sh <node>`
**Topology check:** `bin/netvm-topology.sh` (connectivity-based since 2026-10-04)
**Names:** `bin/netvm-names.sh` — source it: `. bin/netvm-names.sh && netvm_names <node>`
gives you `$NETNS $WG $VETH $VPEER $SUB $GW $PEER_IP`
**DM send:** `python3 bin/dm.py send --agent <from> --to <to> --target main "message"`
**Commit identity** (shared repo — always use per-command flags, never repo config):
```bash
git -c user.name="operator-main" -c user.email="operator-main@frontdoor.local" commit ...
```
**What NOT to do:**
- Don't trust pidfiles for relay health — they lie.
- Don't curl `127.0.0.1:<port>` on the host to check relays — nothing listens there.
- Don't kill a browser that just launched — check `ps -o lstart` first.
- Don't commit other sessions' uncommitted work — use `git add -p` or selective staging.
- Don't restart browsers to fix relay problems — relays and browsers are independent.