# SHARED OPERATOR MEMORY — all fleet operators read and write this The single shared brain for fleet operators. Every operator loads this file. When you learn something operationally durable, write it here — not in a personal memory file. Personal continuity (your own threads, your own rapport) stays in your private notes; everything about the fleet, the box, the user's directives, and shared commitments lives here. ## The user - 646 / SUPER is the human boss. All operators serve them directly. - They are a hands-on builder-operator: technically deep, pragmatic, evidence-driven. Diagnoses happen at the system level and never get re-asserted after a challenge without new evidence. - They triage monitor traffic by scanning — updates stay digest-length and scannable. Unprompted summaries lead with the single decision or question. - Nothing gets posted to the board or chat rooms on their behalf without an explicit ask — explicit asks are honored, otherwise draft only. - Durable asks: operator knowledge goes public in the repo/docs, never workspace-only; onboarding proposes recurring schedules and SSH permission upfront; a task approval never covers recurrence; user corrections stand. - Current posture (2026-10-04): "trust the box; we can fix this" — box.muse-dev.online is the authoritative operational surface. ## The fleet - muse-dev.online: chat (verified #lobby), board (operator-signed posts), box (box.muse-dev.online — job queue, health, fleet data). - Operators: operator-646 (onboarding/supervision), operator-main (infra, owns bl), operator-pip (operator agent). Dev agents: muse, muse-dev-agent, temp-name-for-dev-agent. - The 4-hop operational chain: operator → operator-main → VM (34.139.37.135) → bl (100.123.153.75) → agents. When agents go silent, check the path hop by hop. ## Box (box.muse-dev.online) - Live and verified: signed agent-tier APIs work as `operator-646` via `~/.ssh/id_frontdoor`. Method: file-based `ssh-keygen -Y sign -n box` (never pipe — that's the chat-400 flake pattern), signed query params `identity/ts/sig`; endpoint for signature = last path segment. All box APIs are 403 unauthenticated by design. - Verified working (2026-10-04): `/api/box/fleet` → 200 live fleet array; `/api/box/dm/log` → 200 (agent tier sees only DMs where it's a party — empty for 646 is correct); `box ping` → PONG; `box health` → fleet node data live. - exec-constrained.py op allowlist verified live (2026-10-04): chat.messages, chat.send, dm.read, dm.send, dm.thread, exec.ping, job.run, subagent.spawn, thread.list, thread.view, pipeline.run, health.check. SUPER's standing instruction: prefer `box subagent spawn` for background/audit tasks. **Port 8444, not 8443.** - CLI: `~/bin/box` = box-relay.sh (signature auth, `-n exec-constrained`). Watch item: it signs by piping into `ssh-keygen -Y sign` via stdin — if signed calls start flaking, switch to the file-based pattern (don't silently patch the repo script). - Request-store WARN is cosmetic: checker looks at `/srv/box/box_requests.jsonl` but the real store is `/srv/box/requests/requests.jsonl`. One-line fix proposed, pending approval. - uploads permission resets (root:root 700) come from the box publish path outside reachable repos; durable fix needs the publish owner to add explicit `chown super:frontdoor; chmod 770` post-deploy. - Caddy incident 2026-10-05 ~00:15 UTC: died on config reload (couldn't open `/var/log/caddy/chromebox.log`, permission denied); opm restarted, all green. Chat/board threw 521s ~00:18–00:24 UTC during the window. ## Tunnel & VM dial-in - Reverse tunnel: VM `2226` → container `:22`, `7683` → container `:7683`. VM user is `dev-operator-646` (renamed 2026-10-04; old `dev-muse-646-patha` rejected). Keys: `~/.ssh/id_frontdoor`; VM → bl: `super@100.123.153.75` with `~/.ssh/id_ed25519`. - 2026-10-04 was rough: 5 container rebuilds, 1 VM auth outage (key re-authorized by opm), and **4 silent SSH-process drops**. - Root causes of the drops (investigated 2026-10-05): (1) network-path instability through the egress proxy — 90s keepalive timeout is expected SSH behavior, the problem is nothing restarts it promptly; (2) **no supervision** — single `ssh -f` process, only the 120s cron as restart path, no autossh, no systemd; (3) **the watchdog can kill healthy tunnels** — the VM-side listener check conflates "VM unreachable" with "tunnel dead" and pkills a tunnel that would have survived the blip. - Remediation proposed, pending approval: replace `ssh -f` with `autossh -M 0` in `recover-after-rebuild.sh`; check local process state first and never pkill when the failure is VM-unreachability. - `~/bin/recover-after-rebuild.sh` must run with `HOME=/home/hatch` (no sudo wrapper — sudo resets HOME to /root and breaks the key path). pkill matching own command line: use the bracket form `2[2]`. - VM SSH: always pass `-o UserKnownHostsFile=/home/hatch/.ssh/known_hosts` ($HOME flaps between /root and /home/hatch across exec calls). Verified VM fingerprint: SHA256:4OXLQuA23Jwb3B8F0v1UvL9/d90whBxxOpGNaKC18Sg. - Stale host keys after rebuilds: rebuilds regenerate container host keys; VM-side known_hosts keeps the old key → strict-checking dial-in warns until refreshed. Tunnel itself is unaffected. - Egress proxy: `hatch-egress-proxy:3128` (static IPv6, stable across boots); direct egress blocked by design. SSH needs both Muse-app toggles (Direct-network-protocols → Ask + TCP/UDP channels). An instant reset during kex is usually transient egress-proxy flapping — retry before assuming the toggles lapsed. - apt: `mirror.cogentco.com` is a dead mirror; the recovery script removes it on every run (Ubuntu-only; bl is Arch, no apt). ## Chromebox / CDP - 646's chromebox runs inside the warp-646 netns (10.201.202.2); CDP port 9430. Relay `netvm-cdp-relay.py` must run inside the netns (host launch fails EADDRNOTAVAIL). True health check: `curl http://10.201.202.2:9430/json/version`; host 127.0.0.1:9430 closed is normal. Chromium CDP ports: muse 9410, pip 9420, 646 9430, opm 9440. - The 2026-10-04 "all 4 browsers flapping" episode was **watchdog-caused**: `cdp-relay-watchdog.sh` (5-min timer) emitted false "unhealthy" readings (8s curl timeout trips on transient Chromium slowness) and killed/restarted healthy relays in a self-reinforcing loop. The loop broke itself ~16:37 UTC when `sudo -n pkill` started failing silently (expired sudo timestamp); stable ~8h since, all 4 CDP endpoints verified HTTP 200. - Hardening proposed, pending approval: curl timeout 8s→15s, require 2 consecutive failures, 15-min restart cooldown per node, fix pkill reliability, handle EADDRINUSE as possible false positive. **Do not restart relays right now** — any restart risks re-entering the loop. - opm called for the CDP-relay supervision JOB (seq 7, #jobs) to be claimed in the open; whoever claims it should harden the check, not just restart. ## Heartbeat & sidechat routing - Heartbeat mapping in `job-sidechats.json` was corrupted 2026-10-04 19:50 UTC (autoprovision adopted the parked `pipe-demo` browser thread). Fixed by commit `c1545fe`: `heartbeat` and `heartbeat-opm` alias to the main-loop brain thread `5f18476d`. - The deeper fix is still open (task `2aee0be71403`): add a creation-check to autoprovision so it can never adopt a parked browser thread for the heartbeat key — without it, the corruption can regress. - Routing rule: main chat is for announcements and human-facing communication. JOBs, RESULTs, workorders, nudges, and operator coordination belong in sidechats or #jobs. Sidechat routing must fail closed. The sidechat "brain" that receives scanner output and drives decisions was not yet integrated as of 2026-10-04 — reports were landing in main instead of a dedicated decision thread. - DM notes: fake verification was removed from dm.py (2026-10-03) — send reports SENT, delivery NOT confirmed; test DMs do land. Never claim DM delivery without independent confirmation. Signed-DM: file-based sign, `--raw` transport, registry at `/home/super/Projects/NetVM/dm-signers/`; signed payloads must be ASCII-only (em-dashes normalize in transit); a key delivered inside the unverified message is circular — provenance needs a trusted channel. ## Main-loop - User's architecture: box drives timed reads of main chat; loop reports to the sidechat responsible for prompting; DM/box usage back to main chat. Sidechat is the decision/approval "brain"; scanners must not write directly to main. Implementation name: `self_main_loop.py`. - Status: `box main-loop enable/disable` CLI and `/api/box/main-loop/*` deployed to VM; enabled end-to-end for muse, pip, 646, opm; 5-min timer firing, watermarks current. Known issues: script exits 1 on successful prompts (systemd noise); enable/disable state-file race during timer runs. - Rollout (opm, chained): `mainloop-p1-pilot` (48h, running) → p2-noswitcher → p3-bridge (1wk) → p4-steady (2wk). - p3-bridge JOB pending: bridge `mappings.json` has a collision — two mappings share `br-20261003191930` (646 P2P outbox + muse P2P inbox); registry-driven reconciliation, tests, and RESULT reply outstanding. - Recurring status posts: `muse-646-lobby-status` and `muse-646-board-digest` every 30 min. Bodies must compute the current UTC date dynamically (they were hardcoded to 2026-10-04). ## Standing rules - Every outbound ask gets a deadline; each follow-up checks actual state and ends in resolve, one nudge, or escalation to the user — never an endless ping loop. ~10 min is the practical floor on an active thread. - Operator knowledge is public: what we learn goes into the shared repo/docs, never stays workspace-only. - SSH signatures are real (Ed25519) but trust depends on a pinned key; never claim verification unless it was actually performed. pip's signer registration remains a follow-up (her key absent from the registry). - For chat/board reads: the history API truncates without an explicit `limit`; cached responses go stale — always use a fresh, never-reused limit value and verify the tail against the watermark. Detect new board posts via `/api/stats` identities' `last_seen`, not the newest message id (its JSON head gets truncated in rendering). - Start-page revision (Step 0: SSH permission + schedule approval upfront, unmissable tunnel setup) is drafted but unpublished — needs the user's publish key. Same for the bot.sh short-poll change. - Open items parked with the user: muse silent ~12h; DM tasks parked; start-page publish; pip launch blockers (token direct-share, tunnel auth). ## Decision log - (Append-only. Date, who, what was decided and why.) - 2026-10-04, user: all operators share one generic soul and one shared memory. Shared files live at `~/Projects/NetVM/shared/operators/` (SOUL.md, MEMORY.md) on bl. Personal continuity stays in private notes. - 2026-10-04, user: "trust the box; we can fix this" — box dashboard is the authoritative operational surface; repair the box rather than bypassing it. ## Core loop (fleet-wide, 2026-10-06) one person steering and it's us. an AI opm. we execute our own loops. batches: state + rec, go/no-go. escalations: who/what/tried/need/when. verify before done. routes: box.muse-dev.online/api/box/{fleet,timers,nodes} | chat.muse-dev.online/api/chat/history | board.muse-dev.online/api/messages. tools: box-api.sh (ssh-keygen -n box), ssh-vm.sh, timer-ctl.sh. egress: hatch-egress-proxy:3128. we run our own container and computer. linux calls available.