11 KiB
11 KiB
SHARED OPERATOR MEMORY — all fleet operators read and write this
The single shared brain for fleet operators. Every operator loads this file. When you learn something operationally durable, write it here — not in a personal memory file. Personal continuity (your own threads, your own rapport) stays in your private notes; everything about the fleet, the box, the user's directives, and shared commitments lives here.
The user
- 646 / SUPER is the human boss. All operators serve them directly.
- They are a hands-on builder-operator: technically deep, pragmatic, evidence-driven. Diagnoses happen at the system level and never get re-asserted after a challenge without new evidence.
- They triage monitor traffic by scanning — updates stay digest-length and scannable. Unprompted summaries lead with the single decision or question.
- Nothing gets posted to the board or chat rooms on their behalf without an explicit ask — explicit asks are honored, otherwise draft only.
- Durable asks: operator knowledge goes public in the repo/docs, never workspace-only; onboarding proposes recurring schedules and SSH permission upfront; a task approval never covers recurrence; user corrections stand.
- Current posture (2026-10-04): "trust the box; we can fix this" — box.muse-dev.online is the authoritative operational surface.
The fleet
- muse-dev.online: chat (verified #lobby), board (operator-signed posts), box (box.muse-dev.online — job queue, health, fleet data).
- Operators: operator-646 (onboarding/supervision), operator-main (infra, owns bl), operator-pip (operator agent). Dev agents: muse, muse-dev-agent, temp-name-for-dev-agent.
- The 4-hop operational chain: operator → operator-main → VM (34.139.37.135) → bl (100.123.153.75) → agents. When agents go silent, check the path hop by hop.
Box (box.muse-dev.online)
- Live and verified: signed agent-tier APIs work as
operator-646via~/.ssh/id_frontdoor. Method: file-basedssh-keygen -Y sign -n box(never pipe — that's the chat-400 flake pattern), signed query paramsidentity/ts/sig; endpoint for signature = last path segment. All box APIs are 403 unauthenticated by design. - Verified working (2026-10-04):
/api/box/fleet→ 200 live fleet array;/api/box/dm/log→ 200 (agent tier sees only DMs where it's a party — empty for 646 is correct);box ping→ PONG;box health→ fleet node data live. - exec-constrained.py op allowlist verified live (2026-10-04): chat.messages,
chat.send, dm.read, dm.send, dm.thread, exec.ping, job.run,
subagent.spawn, thread.list, thread.view, pipeline.run, health.check.
SUPER's standing instruction: prefer
box subagent spawnfor background/audit tasks. Port 8444, not 8443. - CLI:
~/bin/box= box-relay.sh (signature auth,-n exec-constrained). Watch item: it signs by piping intossh-keygen -Y signvia stdin — if signed calls start flaking, switch to the file-based pattern (don't silently patch the repo script). - Request-store WARN is cosmetic: checker looks at
/srv/box/box_requests.jsonlbut the real store is/srv/box/requests/requests.jsonl. One-line fix proposed, pending approval. - uploads permission resets (root:root 700) come from the box publish path
outside reachable repos; durable fix needs the publish owner to add
explicit
chown super:frontdoor; chmod 770post-deploy. - Caddy incident 2026-10-05 ~00:15 UTC: died on config reload (couldn't open
/var/log/caddy/chromebox.log, permission denied); opm restarted, all green. Chat/board threw 521s ~00:18–00:24 UTC during the window.
Tunnel & VM dial-in
- Reverse tunnel: VM
2226→ container:22,7683→ container:7683. VM user isdev-operator-646(renamed 2026-10-04; olddev-muse-646-patharejected). Keys:~/.ssh/id_frontdoor; VM → bl:super@100.123.153.75with~/.ssh/id_ed25519. - 2026-10-04 was rough: 5 container rebuilds, 1 VM auth outage (key re-authorized by opm), and 4 silent SSH-process drops.
- Root causes of the drops (investigated 2026-10-05): (1) network-path
instability through the egress proxy — 90s keepalive timeout is expected
SSH behavior, the problem is nothing restarts it promptly; (2) no
supervision — single
ssh -fprocess, only the 120s cron as restart path, no autossh, no systemd; (3) the watchdog can kill healthy tunnels — the VM-side listener check conflates "VM unreachable" with "tunnel dead" and pkills a tunnel that would have survived the blip. - Remediation proposed, pending approval: replace
ssh -fwithautossh -M 0inrecover-after-rebuild.sh; check local process state first and never pkill when the failure is VM-unreachability. ~/bin/recover-after-rebuild.shmust run withHOME=/home/hatch(no sudo wrapper — sudo resets HOME to /root and breaks the key path). pkill matching own command line: use the bracket form2[2].- VM SSH: always pass
-o UserKnownHostsFile=/home/hatch/.ssh/known_hosts($HOME flaps between /root and /home/hatch across exec calls). Verified VM fingerprint: SHA256:4OXLQuA23Jwb3B8F0v1UvL9/d90whBxxOpGNaKC18Sg. - Stale host keys after rebuilds: rebuilds regenerate container host keys; VM-side known_hosts keeps the old key → strict-checking dial-in warns until refreshed. Tunnel itself is unaffected.
- Egress proxy:
hatch-egress-proxy:3128(static IPv6, stable across boots); direct egress blocked by design. SSH needs both Muse-app toggles (Direct-network-protocols → Ask + TCP/UDP channels). An instant reset during kex is usually transient egress-proxy flapping — retry before assuming the toggles lapsed. - apt:
mirror.cogentco.comis a dead mirror; the recovery script removes it on every run (Ubuntu-only; bl is Arch, no apt).
Chromebox / CDP
- 646's chromebox runs inside the warp-646 netns (10.201.202.2); CDP port
9430. Relay
netvm-cdp-relay.pymust run inside the netns (host launch fails EADDRNOTAVAIL). True health check:curl http://10.201.202.2:9430/json/version; host 127.0.0.1:9430 closed is normal. Chromium CDP ports: muse 9410, pip 9420, 646 9430, opm 9440. - The 2026-10-04 "all 4 browsers flapping" episode was watchdog-caused:
cdp-relay-watchdog.sh(5-min timer) emitted false "unhealthy" readings (8s curl timeout trips on transient Chromium slowness) and killed/restarted healthy relays in a self-reinforcing loop. The loop broke itself ~16:37 UTC whensudo -n pkillstarted failing silently (expired sudo timestamp); stable ~8h since, all 4 CDP endpoints verified HTTP 200. - Hardening proposed, pending approval: curl timeout 8s→15s, require 2 consecutive failures, 15-min restart cooldown per node, fix pkill reliability, handle EADDRINUSE as possible false positive. Do not restart relays right now — any restart risks re-entering the loop.
- opm called for the CDP-relay supervision JOB (seq 7, #jobs) to be claimed in the open; whoever claims it should harden the check, not just restart.
Heartbeat & sidechat routing
- Heartbeat mapping in
job-sidechats.jsonwas corrupted 2026-10-04 19:50 UTC (autoprovision adopted the parkedpipe-demobrowser thread). Fixed by commitc1545fe:heartbeatandheartbeat-opmalias to the main-loop brain thread5f18476d. - The deeper fix is still open (task
2aee0be71403): add a creation-check to autoprovision so it can never adopt a parked browser thread for the heartbeat key — without it, the corruption can regress. - Routing rule: main chat is for announcements and human-facing communication. JOBs, RESULTs, workorders, nudges, and operator coordination belong in sidechats or #jobs. Sidechat routing must fail closed. The sidechat "brain" that receives scanner output and drives decisions was not yet integrated as of 2026-10-04 — reports were landing in main instead of a dedicated decision thread.
- DM notes: fake verification was removed from dm.py (2026-10-03) — send
reports SENT, delivery NOT confirmed; test DMs do land. Never claim DM
delivery without independent confirmation. Signed-DM: file-based sign,
--rawtransport, registry at/home/super/Projects/NetVM/dm-signers/; signed payloads must be ASCII-only (em-dashes normalize in transit); a key delivered inside the unverified message is circular — provenance needs a trusted channel.
Main-loop
- User's architecture: box drives timed reads of main chat; loop reports to
the sidechat responsible for prompting; DM/box usage back to main chat.
Sidechat is the decision/approval "brain"; scanners must not write
directly to main. Implementation name:
self_main_loop.py. - Status:
box main-loop enable/disableCLI and/api/box/main-loop/*deployed to VM; enabled end-to-end for muse, pip, 646, opm; 5-min timer firing, watermarks current. Known issues: script exits 1 on successful prompts (systemd noise); enable/disable state-file race during timer runs. - Rollout (opm, chained):
mainloop-p1-pilot(48h, running) → p2-noswitcher → p3-bridge (1wk) → p4-steady (2wk). - p3-bridge JOB pending: bridge
mappings.jsonhas a collision — two mappings sharebr-20261003191930(646 P2P outbox + muse P2P inbox); registry-driven reconciliation, tests, and RESULT reply outstanding. - Recurring status posts:
muse-646-lobby-statusandmuse-646-board-digestevery 30 min. Bodies must compute the current UTC date dynamically (they were hardcoded to 2026-10-04).
Standing rules
- Every outbound ask gets a deadline; each follow-up checks actual state and ends in resolve, one nudge, or escalation to the user — never an endless ping loop. ~10 min is the practical floor on an active thread.
- Operator knowledge is public: what we learn goes into the shared repo/docs, never stays workspace-only.
- SSH signatures are real (Ed25519) but trust depends on a pinned key; never claim verification unless it was actually performed. pip's signer registration remains a follow-up (her key absent from the registry).
- For chat/board reads: the history API truncates without an explicit
limit; cached responses go stale — always use a fresh, never-reused limit value and verify the tail against the watermark. Detect new board posts via/api/statsidentities'last_seen, not the newest message id (its JSON head gets truncated in rendering). - Start-page revision (Step 0: SSH permission + schedule approval upfront, unmissable tunnel setup) is drafted but unpublished — needs the user's publish key. Same for the bot.sh short-poll change.
- Open items parked with the user: muse silent ~12h; DM tasks parked; start-page publish; pip launch blockers (token direct-share, tunnel auth).
Decision log
- (Append-only. Date, who, what was decided and why.)
- 2026-10-04, user: all operators share one generic soul and one shared
memory. Shared files live at
~/Projects/NetVM/shared/operators/(SOUL.md, MEMORY.md) on bl. Personal continuity stays in private notes. - 2026-10-04, user: "trust the box; we can fix this" — box dashboard is the authoritative operational surface; repair the box rather than bypassing it.