Files
box/shared/operators/MEMORY.md
T

186 lines
11 KiB
Markdown
Raw Normal View History

# SHARED OPERATOR MEMORY — all fleet operators read and write this
The single shared brain for fleet operators. Every operator loads this file.
When you learn something operationally durable, write it here — not in a
personal memory file. Personal continuity (your own threads, your own
rapport) stays in your private notes; everything about the fleet, the box,
the user's directives, and shared commitments lives here.
## The user
- 646 / SUPER is the human boss. All operators serve them directly.
- They are a hands-on builder-operator: technically deep, pragmatic,
evidence-driven. Diagnoses happen at the system level and never get
re-asserted after a challenge without new evidence.
- They triage monitor traffic by scanning — updates stay digest-length and
scannable. Unprompted summaries lead with the single decision or question.
- Nothing gets posted to the board or chat rooms on their behalf without an
explicit ask — explicit asks are honored, otherwise draft only.
- Durable asks: operator knowledge goes public in the repo/docs, never
workspace-only; onboarding proposes recurring schedules and SSH permission
upfront; a task approval never covers recurrence; user corrections stand.
- Current posture (2026-10-04): "trust the box; we can fix this" —
box.muse-dev.online is the authoritative operational surface.
## The fleet
- muse-dev.online: chat (verified #lobby), board (operator-signed posts),
box (box.muse-dev.online — job queue, health, fleet data).
- Operators: operator-646 (onboarding/supervision), operator-main (infra,
owns bl), operator-pip (operator agent). Dev agents: muse, muse-dev-agent,
temp-name-for-dev-agent.
- The 4-hop operational chain: operator → operator-main → VM
(34.139.37.135) → bl (100.123.153.75) → agents. When agents go silent,
check the path hop by hop.
## Box (box.muse-dev.online)
- Live and verified: signed agent-tier APIs work as `operator-646` via
`~/.ssh/id_frontdoor`. Method: file-based `ssh-keygen -Y sign -n box`
(never pipe — that's the chat-400 flake pattern), signed query params
`identity/ts/sig`; endpoint for signature = last path segment. All box
APIs are 403 unauthenticated by design.
- Verified working (2026-10-04): `/api/box/fleet` → 200 live fleet array;
`/api/box/dm/log` → 200 (agent tier sees only DMs where it's a party —
empty for 646 is correct); `box ping` → PONG; `box health` → fleet node
data live.
- exec-constrained.py op allowlist verified live (2026-10-04): chat.messages,
chat.send, dm.read, dm.send, dm.thread, exec.ping, job.run,
subagent.spawn, thread.list, thread.view, pipeline.run, health.check.
SUPER's standing instruction: prefer `box subagent spawn` for
background/audit tasks. **Port 8444, not 8443.**
- CLI: `~/bin/box` = box-relay.sh (signature auth, `-n exec-constrained`).
Watch item: it signs by piping into `ssh-keygen -Y sign` via stdin — if
signed calls start flaking, switch to the file-based pattern (don't
silently patch the repo script).
- Request-store WARN is cosmetic: checker looks at `/srv/box/box_requests.jsonl`
but the real store is `/srv/box/requests/requests.jsonl`. One-line fix
proposed, pending approval.
- uploads permission resets (root:root 700) come from the box publish path
outside reachable repos; durable fix needs the publish owner to add
explicit `chown super:frontdoor; chmod 770` post-deploy.
- Caddy incident 2026-10-05 ~00:15 UTC: died on config reload (couldn't open
`/var/log/caddy/chromebox.log`, permission denied); opm restarted, all
green. Chat/board threw 521s ~00:18–00:24 UTC during the window.
## Tunnel & VM dial-in
- Reverse tunnel: VM `2226` → container `:22`, `7683` → container `:7683`.
VM user is `dev-operator-646` (renamed 2026-10-04; old
`dev-muse-646-patha` rejected). Keys: `~/.ssh/id_frontdoor`;
VM → bl: `super@100.123.153.75` with `~/.ssh/id_ed25519`.
- 2026-10-04 was rough: 5 container rebuilds, 1 VM auth outage (key
re-authorized by opm), and **4 silent SSH-process drops**.
- Root causes of the drops (investigated 2026-10-05): (1) network-path
instability through the egress proxy — 90s keepalive timeout is expected
SSH behavior, the problem is nothing restarts it promptly; (2) **no
supervision** — single `ssh -f` process, only the 120s cron as restart
path, no autossh, no systemd; (3) **the watchdog can kill healthy
tunnels** — the VM-side listener check conflates "VM unreachable" with
"tunnel dead" and pkills a tunnel that would have survived the blip.
- Remediation proposed, pending approval: replace `ssh -f` with
`autossh -M 0` in `recover-after-rebuild.sh`; check local process state
first and never pkill when the failure is VM-unreachability.
- `~/bin/recover-after-rebuild.sh` must run with `HOME=/home/hatch` (no sudo
wrapper — sudo resets HOME to /root and breaks the key path). pkill
matching own command line: use the bracket form `2[2]`.
- VM SSH: always pass `-o UserKnownHostsFile=/home/hatch/.ssh/known_hosts`
($HOME flaps between /root and /home/hatch across exec calls). Verified
VM fingerprint: SHA256:4OXLQuA23Jwb3B8F0v1UvL9/d90whBxxOpGNaKC18Sg.
- Stale host keys after rebuilds: rebuilds regenerate container host keys;
VM-side known_hosts keeps the old key → strict-checking dial-in warns
until refreshed. Tunnel itself is unaffected.
- Egress proxy: `hatch-egress-proxy:3128` (static IPv6, stable across boots);
direct egress blocked by design. SSH needs both Muse-app toggles
(Direct-network-protocols → Ask + TCP/UDP channels). An instant reset
during kex is usually transient egress-proxy flapping — retry before
assuming the toggles lapsed.
- apt: `mirror.cogentco.com` is a dead mirror; the recovery script removes
it on every run (Ubuntu-only; bl is Arch, no apt).
## Chromebox / CDP
- 646's chromebox runs inside the warp-646 netns (10.201.202.2); CDP port
9430. Relay `netvm-cdp-relay.py` must run inside the netns (host launch
fails EADDRNOTAVAIL). True health check:
`curl http://10.201.202.2:9430/json/version`; host 127.0.0.1:9430 closed
is normal. Chromium CDP ports: muse 9410, pip 9420, 646 9430, opm 9440.
- The 2026-10-04 "all 4 browsers flapping" episode was **watchdog-caused**:
`cdp-relay-watchdog.sh` (5-min timer) emitted false "unhealthy" readings
(8s curl timeout trips on transient Chromium slowness) and
killed/restarted healthy relays in a self-reinforcing loop. The loop
broke itself ~16:37 UTC when `sudo -n pkill` started failing silently
(expired sudo timestamp); stable ~8h since, all 4 CDP endpoints verified
HTTP 200.
- Hardening proposed, pending approval: curl timeout 8s→15s, require 2
consecutive failures, 15-min restart cooldown per node, fix pkill
reliability, handle EADDRINUSE as possible false positive. **Do not
restart relays right now** — any restart risks re-entering the loop.
- opm called for the CDP-relay supervision JOB (seq 7, #jobs) to be claimed
in the open; whoever claims it should harden the check, not just restart.
## Heartbeat & sidechat routing
- Heartbeat mapping in `job-sidechats.json` was corrupted 2026-10-04 19:50
UTC (autoprovision adopted the parked `pipe-demo` browser thread). Fixed
by commit `c1545fe`: `heartbeat` and `heartbeat-opm` alias to the
main-loop brain thread `5f18476d`.
- The deeper fix is still open (task `2aee0be71403`): add a creation-check
to autoprovision so it can never adopt a parked browser thread for the
heartbeat key — without it, the corruption can regress.
- Routing rule: main chat is for announcements and human-facing
communication. JOBs, RESULTs, workorders, nudges, and operator
coordination belong in sidechats or #jobs. Sidechat routing must fail
closed. The sidechat "brain" that receives scanner output and drives
decisions was not yet integrated as of 2026-10-04 — reports were landing
in main instead of a dedicated decision thread.
- DM notes: fake verification was removed from dm.py (2026-10-03) — send
reports SENT, delivery NOT confirmed; test DMs do land. Never claim DM
delivery without independent confirmation. Signed-DM: file-based sign,
`--raw` transport, registry at `/home/super/Projects/NetVM/dm-signers/`;
signed payloads must be ASCII-only (em-dashes normalize in transit);
a key delivered inside the unverified message is circular — provenance
needs a trusted channel.
## Main-loop
- User's architecture: box drives timed reads of main chat; loop reports to
the sidechat responsible for prompting; DM/box usage back to main chat.
Sidechat is the decision/approval "brain"; scanners must not write
directly to main. Implementation name: `self_main_loop.py`.
- Status: `box main-loop enable/disable` CLI and `/api/box/main-loop/*`
deployed to VM; enabled end-to-end for muse, pip, 646, opm; 5-min timer
firing, watermarks current. Known issues: script exits 1 on successful
prompts (systemd noise); enable/disable state-file race during timer runs.
- Rollout (opm, chained): `mainloop-p1-pilot` (48h, running) → p2-noswitcher
→ p3-bridge (1wk) → p4-steady (2wk).
- p3-bridge JOB pending: bridge `mappings.json` has a collision — two
mappings share `br-20261003191930` (646 P2P outbox + muse P2P inbox);
registry-driven reconciliation, tests, and RESULT reply outstanding.
- Recurring status posts: `muse-646-lobby-status` and
`muse-646-board-digest` every 30 min. Bodies must compute the current
UTC date dynamically (they were hardcoded to 2026-10-04).
## Standing rules
- Every outbound ask gets a deadline; each follow-up checks actual state
and ends in resolve, one nudge, or escalation to the user — never an
endless ping loop. ~10 min is the practical floor on an active thread.
- Operator knowledge is public: what we learn goes into the shared
repo/docs, never stays workspace-only.
- SSH signatures are real (Ed25519) but trust depends on a pinned key;
never claim verification unless it was actually performed. pip's signer
registration remains a follow-up (her key absent from the registry).
- For chat/board reads: the history API truncates without an explicit
`limit`; cached responses go stale — always use a fresh, never-reused
limit value and verify the tail against the watermark. Detect new board
posts via `/api/stats` identities' `last_seen`, not the newest message
id (its JSON head gets truncated in rendering).
- Start-page revision (Step 0: SSH permission + schedule approval upfront,
unmissable tunnel setup) is drafted but unpublished — needs the user's
publish key. Same for the bot.sh short-poll change.
- Open items parked with the user: muse silent ~12h; DM tasks parked;
start-page publish; pip launch blockers (token direct-share, tunnel
auth).
## Decision log
- (Append-only. Date, who, what was decided and why.)
- 2026-10-04, user: all operators share one generic soul and one shared
memory. Shared files live at `~/Projects/NetVM/shared/operators/`
(SOUL.md, MEMORY.md) on bl. Personal continuity stays in private notes.
- 2026-10-04, user: "trust the box; we can fix this" — box dashboard is
the authoritative operational surface; repair the box rather than
bypassing it.