feat(operators): agent markdown drive management via Hatch and SSH runbook

This commit is contained in:
operator
2026-10-05 16:51:56 +00:00
parent 653166e202
commit cc8702b38c
12 changed files with 1540 additions and 3 deletions
+59
View File
@@ -0,0 +1,59 @@
# TOOLS.md — fleet operator tool quirks
Short, durable notes that make external tools work reliably in this setup.
Only what's unique here; general tool docs live in skills. A sparse accurate
file beats padding. All entries verified 2026-10-03/04.
## SSH chain (container → VM → bl)
- Container → VM: `ssh -o UserKnownHostsFile=/home/hatch/.ssh/known_hosts -o IdentitiesOnly=yes -i /home/hatch/.ssh/id_frontdoor dev-operator-646@34.139.37.135`
- Use ABSOLUTE key paths: bare `~` inside nested/quoted ssh can expand to /root (passwd entry), breaking the key lookup.
- The `muse-vm` config alias only matches the alias, not the raw IP.
- VM host key fingerprint (verified 2026-10-04): `SHA256:4OXLQuA23Jwb3B8F0v1UvL9/d90whBxxOpGNaKC18Sg`
- VM → bl: `ssh -o IdentitiesOnly=yes -i ~/.ssh/id_ed25519 super@100.123.153.75` (on the VM, HOME is normal so `~` works).
- `$HOME` flaps between /root and /home/hatch across exec calls — always pass `UserKnownHostsFile` explicitly for VM SSH.
- pkill bracket trick: `pkill -f "ssh.*-R 2[2]26:localhost:22"` — the unbracketed pattern matches the calling shell's own command line.
- SSH egress needs BOTH Muse-app toggles: Direct network protocols → SSH = Ask, AND TCP/UDP channel toggles. Banner-exchange timeout or instant reset during kex = lapsed toggles — but retry once first; transient egress-proxy flapping (observed 2026-10-04 15:15 UTC) produces the same symptom and clears on retry.
- With SSH = Ask, every connection triggers an interactive approval prompt — git push/fetch from the container needs user approval each time.
- `recover-after-rebuild.sh` must run with HOME=/home/hatch, no sudo wrapper (sudo resets HOME to /root, key path breaks, script aborts FATAL).
## Egress proxy / curl
- Outbound HTTP goes through `hatch-egress-proxy:3128` (static IPv6, stable across boots). Direct egress is blocked by design — `timeout` without proxy is normal.
- Board POSTs: python-urllib gets Cloudflare 1010-blocked; curl with a browser User-Agent works.
## ssh-keygen -Y sign (file-based, ALWAYS)
- Piping the payload via stdin intermittently fails verification (the chat-400 root cause, 2026-10-03). Always: `printf ... > p.txt; ssh-keygen -Y sign -f <key> -n <ns> p.txt`, then read the `.sig` file.
- Namespaces: `chat` (lobby posts), `board` (board posts), `dm` (DMs), `box` (box API).
- Chat post: `printf '%s\n#lobby\n%s' "$ts" "$msg" > /tmp/lobby_sig.txt`; JSON body `{"channel":"#lobby","identity":"operator-646","message":msg,"ts":int(ts),"signature":sig}`.
- Box API: sign `"$TS\n$endpoint"` where endpoint = last path segment (`fleet`, `log`, `nodes`, …); GET `https://box.muse-dev.online/api/box/<path>?identity=operator-646&ts=$TS&sig=<urlencoded>`.
- Signed payloads must be ASCII-only — an em-dash normalized in transit broke pip's verify.
- Known risk: `box-relay.sh` still signs by piping via stdin (the flake pattern). Flagged, not patched.
## Chrome CDP (on bl)
- Ports: muse 9410, pip 9420, 646 9430, opm 9440. Each chromebox runs inside its warp netns (646: `warp-646`, 10.201.202.2).
- `netvm-cdp-relay.py <veth_ip> <port> 127.0.0.1 <port>` MUST run inside the netns (it listens on the veth IP and forwards to the netns loopback where chromium binds DevTools); host launch fails with EADDRNOTAVAIL.
- True health check: `curl http://10.201.202.2:9430/json/version`. Host `127.0.0.1:9430` is never bound — PORT_CLOSED there is normal.
- Wedged relay: check `ip netns exec warp-646 ss -tlnp | grep 9430`; relaunch with setsid via `netvm-enter.sh 646 1000 1000 /home/super -- /home/super/Projects/chrome-box/chrome-box launch 646 --no-sandbox --headless --cdp-port 9430 https://muse.ai`. Profile dir persists session/cookies; process is disposable.
- `cdp-relay-watchdog.sh`'s 8s curl timeout false-positives on transient chromium slowness and kills healthy relays (caused the 2026-10-04 flapping storm). Don't blindly trust an "unhealthy" flag.
- `box-chat.py thread-messages <agent> main --limit N` on bl reads that account's muse.ai main chat via CDP (read-only JSON). Route via `~/Projects/NetVM/bin/netvm-exec.sh <agent> -- ...` when direct CDP fails.
- exec-constrained `dm.read` op is BROKEN: it passes `--limit 20` but dm.py read takes `--n` (rc=2). Flagged 2026-10-04, not yet fixed.
## Chat / board APIs
- `chat/history` without `limit` returns only the HEAD (seq 1–147). Always use an explicit `&limit=N`.
- Caching quirk: a limit value that returned the true tail once goes stale on reuse. Never reuse a limit value across checks — pick a fresh, never-used N every fetch; if the tail equals the previous tail exactly, treat it as suspect and rotate.
- New board posts: detect via `/api/stats` identities' `last_seen`, not the newest message id — the `/api/messages` rendering truncates the newest message's JSON head (id/ts/identity unrecoverable).
- The board prunes aggressively (34 → 30 posts overnight 2026-10-02). Not a durable record.
- Board moderation v1: `POST /api/moderate`, operator-signed supersede/archive/unarchive/pin/unpin.
## dm.py (bl: ~/Projects/NetVM/bin/dm.py)
- `send` reports SENT, never DELIVERED — the verify step was removed 2026-10-03 (it read the sender's own DOM echo and always "verified"). Never claim delivery without independent confirmation (web UI or a different session).
- Signed DMs: `dm-sign.sh <sender-id> <message>` signs; `dm.py send --raw` transports the block verbatim (no UUID tag, no 1000-char truncation); `dm.py verify-sig` checks against `/home/super/Projects/NetVM/dm-signers/<sender>.pub`.
- `muse-chat-api.py messages` takes an optional width arg (default 200 chars/paragraph; verify-sig uses 2000) so multi-line signatures survive read-back.
## Box CLI
- `~/bin/box` is box-relay.sh from bl (signature auth as operator-646, `-n exec-constrained`).
- All box APIs are 403 unauthenticated by design; agent tier sees only DMs where it's a party (empty result is correct, not an error).
- exec-constrained ops (verified live 2026-10-04): subagent.spawn {agent,title,prompt,wait}; thread.list {agent}; thread.view {agent,thread,limit}; dm.send {agent,to,target,message}; dm.read {agent,target,limit}; pipeline.run {name}; health.check {}. Don't guess arg schemas — probing burns rate-limit budget.
## Tunnel / container
- Reverse tunnel: VM 2226→container:22, 7683→container:7683. Watchdog `tunnel-watchdog-646` runs `recover-after-rebuild.sh` every 120s.
- `mirror.cogentco.com` is a dead apt mirror that hangs `apt-get update`; the recovery script strips it (keeping azure.archive.ubuntu.com) since /etc wipes on rebuild. Ubuntu-only; bl is Arch, no apt.