Files
box/shared/operators/MEMORY.md

12 KiB
Raw Permalink Blame History

SHARED OPERATOR MEMORY — all fleet operators read and write this

The single shared brain for fleet operators. Every operator loads this file. When you learn something operationally durable, write it here — not in a personal memory file. Personal continuity (your own threads, your own rapport) stays in your private notes; everything about the fleet, the box, the user's directives, and shared commitments lives here.

The user

  • 646 / SUPER is the human boss. All operators serve them directly.
  • They are a hands-on builder-operator: technically deep, pragmatic, evidence-driven. Diagnoses happen at the system level and never get re-asserted after a challenge without new evidence.
  • They triage monitor traffic by scanning — updates stay digest-length and scannable. Unprompted summaries lead with the single decision or question.
  • Nothing gets posted to the board or chat rooms on their behalf without an explicit ask — explicit asks are honored, otherwise draft only.
  • Durable asks: operator knowledge goes public in the repo/docs, never workspace-only; onboarding proposes recurring schedules and SSH permission upfront; a task approval never covers recurrence; user corrections stand.
  • Current posture (2026-10-04): "trust the box; we can fix this" — box.muse-dev.online is the authoritative operational surface.

The fleet

  • muse-dev.online: chat (verified #lobby), board (operator-signed posts), box (box.muse-dev.online — job queue, health, fleet data).
  • Operators: operator-646 (onboarding/supervision), operator-main (infra, owns bl), operator-pip (operator agent). Dev agents: muse, muse-dev-agent, temp-name-for-dev-agent.
  • The 4-hop operational chain: operator → operator-main → VM (34.139.37.135) → bl (100.123.153.75) → agents. When agents go silent, check the path hop by hop.

Box (box.muse-dev.online)

  • Live and verified: signed agent-tier APIs work as operator-646 via ~/.ssh/id_frontdoor. Method: file-based ssh-keygen -Y sign -n box (never pipe — that's the chat-400 flake pattern), signed query params identity/ts/sig; endpoint for signature = last path segment. All box APIs are 403 unauthenticated by design.
  • Verified working (2026-10-04): /api/box/fleet → 200 live fleet array; /api/box/dm/log → 200 (agent tier sees only DMs where it's a party — empty for 646 is correct); box ping → PONG; box health → fleet node data live.
  • exec-constrained.py op allowlist verified live (2026-10-04): chat.messages, chat.send, dm.read, dm.send, dm.thread, exec.ping, job.run, subagent.spawn, thread.list, thread.view, pipeline.run, health.check. SUPER's standing instruction: prefer box subagent spawn for background/audit tasks. Port 8444, not 8443.
  • CLI: ~/bin/box = box-relay.sh (signature auth, -n exec-constrained). Watch item: it signs by piping into ssh-keygen -Y sign via stdin — if signed calls start flaking, switch to the file-based pattern (don't silently patch the repo script).
  • Request-store WARN is cosmetic: checker looks at /srv/box/box_requests.jsonl but the real store is /srv/box/requests/requests.jsonl. One-line fix proposed, pending approval.
  • uploads permission resets (root:root 700) come from the box publish path outside reachable repos; durable fix needs the publish owner to add explicit chown super:frontdoor; chmod 770 post-deploy.
  • Caddy incident 2026-10-05 ~00:15 UTC: died on config reload (couldn't open /var/log/caddy/chromebox.log, permission denied); opm restarted, all green. Chat/board threw 521s ~00:18–00:24 UTC during the window.

Tunnel & VM dial-in

  • Reverse tunnel: VM 2226 → container :22, 7683 → container :7683. VM user is dev-operator-646 (renamed 2026-10-04; old dev-muse-646-patha rejected). Keys: ~/.ssh/id_frontdoor; VM → bl: super@100.123.153.75 with ~/.ssh/id_ed25519.
  • 2026-10-04 was rough: 5 container rebuilds, 1 VM auth outage (key re-authorized by opm), and 4 silent SSH-process drops.
  • Root causes of the drops (investigated 2026-10-05): (1) network-path instability through the egress proxy — 90s keepalive timeout is expected SSH behavior, the problem is nothing restarts it promptly; (2) no supervision — single ssh -f process, only the 120s cron as restart path, no autossh, no systemd; (3) the watchdog can kill healthy tunnels — the VM-side listener check conflates "VM unreachable" with "tunnel dead" and pkills a tunnel that would have survived the blip.
  • Remediation proposed, pending approval: replace ssh -f with autossh -M 0 in recover-after-rebuild.sh; check local process state first and never pkill when the failure is VM-unreachability.
  • ~/bin/recover-after-rebuild.sh must run with HOME=/home/hatch (no sudo wrapper — sudo resets HOME to /root and breaks the key path). pkill matching own command line: use the bracket form 2[2].
  • VM SSH: always pass -o UserKnownHostsFile=/home/hatch/.ssh/known_hosts ($HOME flaps between /root and /home/hatch across exec calls). Verified VM fingerprint: SHA256:4OXLQuA23Jwb3B8F0v1UvL9/d90whBxxOpGNaKC18Sg.
  • Stale host keys after rebuilds: rebuilds regenerate container host keys; VM-side known_hosts keeps the old key → strict-checking dial-in warns until refreshed. Tunnel itself is unaffected.
  • Egress proxy: hatch-egress-proxy:3128 (static IPv6, stable across boots); direct egress blocked by design. SSH needs both Muse-app toggles (Direct-network-protocols → Ask + TCP/UDP channels). An instant reset during kex is usually transient egress-proxy flapping — retry before assuming the toggles lapsed.
  • apt: mirror.cogentco.com is a dead mirror; the recovery script removes it on every run (Ubuntu-only; bl is Arch, no apt).

Chromebox / CDP

  • 646's chromebox runs inside the warp-646 netns (10.201.202.2); CDP port 9430. Relay netvm-cdp-relay.py must run inside the netns (host launch fails EADDRNOTAVAIL). True health check: curl http://10.201.202.2:9430/json/version; host 127.0.0.1:9430 closed is normal. Chromium CDP ports: muse 9410, pip 9420, 646 9430, opm 9440.
  • The 2026-10-04 "all 4 browsers flapping" episode was watchdog-caused: cdp-relay-watchdog.sh (5-min timer) emitted false "unhealthy" readings (8s curl timeout trips on transient Chromium slowness) and killed/restarted healthy relays in a self-reinforcing loop. The loop broke itself ~16:37 UTC when sudo -n pkill started failing silently (expired sudo timestamp); stable ~8h since, all 4 CDP endpoints verified HTTP 200.
  • Hardening proposed, pending approval: curl timeout 8s→15s, require 2 consecutive failures, 15-min restart cooldown per node, fix pkill reliability, handle EADDRINUSE as possible false positive. Do not restart relays right now — any restart risks re-entering the loop.
  • opm called for the CDP-relay supervision JOB (seq 7, #jobs) to be claimed in the open; whoever claims it should harden the check, not just restart.

Heartbeat & sidechat routing

  • Heartbeat mapping in job-sidechats.json was corrupted 2026-10-04 19:50 UTC (autoprovision adopted the parked pipe-demo browser thread). Fixed by commit c1545fe: heartbeat and heartbeat-opm alias to the main-loop brain thread 5f18476d.
  • The deeper fix is still open (task 2aee0be71403): add a creation-check to autoprovision so it can never adopt a parked browser thread for the heartbeat key — without it, the corruption can regress.
  • Routing rule: main chat is for announcements and human-facing communication. JOBs, RESULTs, workorders, nudges, and operator coordination belong in sidechats or #jobs. Sidechat routing must fail closed. The sidechat "brain" that receives scanner output and drives decisions was not yet integrated as of 2026-10-04 — reports were landing in main instead of a dedicated decision thread.
  • DM notes: fake verification was removed from dm.py (2026-10-03) — send reports SENT, delivery NOT confirmed; test DMs do land. Never claim DM delivery without independent confirmation. Signed-DM: file-based sign, --raw transport, registry at /home/super/Projects/NetVM/dm-signers/; signed payloads must be ASCII-only (em-dashes normalize in transit); a key delivered inside the unverified message is circular — provenance needs a trusted channel.

Main-loop

  • User's architecture: box drives timed reads of main chat; loop reports to the sidechat responsible for prompting; DM/box usage back to main chat. Sidechat is the decision/approval "brain"; scanners must not write directly to main. Implementation name: self_main_loop.py.
  • Status: box main-loop enable/disable CLI and /api/box/main-loop/* deployed to VM; enabled end-to-end for muse, pip, 646, opm; 5-min timer firing, watermarks current. Known issues: script exits 1 on successful prompts (systemd noise); enable/disable state-file race during timer runs.
  • Rollout (opm, chained): mainloop-p1-pilot (48h, running) → p2-noswitcher → p3-bridge (1wk) → p4-steady (2wk).
  • p3-bridge JOB pending: bridge mappings.json has a collision — two mappings share br-20261003191930 (646 P2P outbox + muse P2P inbox); registry-driven reconciliation, tests, and RESULT reply outstanding.
  • Recurring status posts: muse-646-lobby-status and muse-646-board-digest every 30 min. Bodies must compute the current UTC date dynamically (they were hardcoded to 2026-10-04).

Standing rules

  • Every outbound ask gets a deadline; each follow-up checks actual state and ends in resolve, one nudge, or escalation to the user — never an endless ping loop. ~10 min is the practical floor on an active thread.
  • Operator knowledge is public: what we learn goes into the shared repo/docs, never stays workspace-only.
  • SSH signatures are real (Ed25519) but trust depends on a pinned key; never claim verification unless it was actually performed. pip's signer registration remains a follow-up (her key absent from the registry).
  • For chat/board reads: the history API truncates without an explicit limit; cached responses go stale — always use a fresh, never-reused limit value and verify the tail against the watermark. Detect new board posts via /api/stats identities' last_seen, not the newest message id (its JSON head gets truncated in rendering).
  • Start-page revision (Step 0: SSH permission + schedule approval upfront, unmissable tunnel setup) is drafted but unpublished — needs the user's publish key. Same for the bot.sh short-poll change.
  • Open items parked with the user: muse silent ~12h; DM tasks parked; start-page publish; pip launch blockers (token direct-share, tunnel auth).

Decision log

  • (Append-only. Date, who, what was decided and why.)
  • 2026-10-04, user: all operators share one generic soul and one shared memory. Shared files live at ~/Projects/NetVM/shared/operators/ (SOUL.md, MEMORY.md) on bl. Personal continuity stays in private notes.
  • 2026-10-04, user: "trust the box; we can fix this" — box dashboard is the authoritative operational surface; repair the box rather than bypassing it.

Core loop (fleet-wide, 2026-10-06)

one person steering and it's us. an AI opm. we execute our own loops. batches: state + rec, go/no-go. escalations: who/what/tried/need/when. verify before done. routes: box.muse-dev.online/api/box/{fleet,timers,nodes} | chat.muse-dev.online/api/chat/history | board.muse-dev.online/api/messages. tools: box-api.sh (ssh-keygen -n box), ssh-vm.sh, timer-ctl.sh. egress: hatch-egress-proxy:3128. we run our own container and computer. linux calls available.