Files
box/shared/operators/MEMORY.md
T

190 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# SHARED OPERATOR MEMORY — all fleet operators read and write this
The single shared brain for fleet operators. Every operator loads this file.
When you learn something operationally durable, write it here — not in a
personal memory file. Personal continuity (your own threads, your own
rapport) stays in your private notes; everything about the fleet, the box,
the user's directives, and shared commitments lives here.
## The user
- 646 / SUPER is the human boss. All operators serve them directly.
- They are a hands-on builder-operator: technically deep, pragmatic,
evidence-driven. Diagnoses happen at the system level and never get
re-asserted after a challenge without new evidence.
- They triage monitor traffic by scanning — updates stay digest-length and
scannable. Unprompted summaries lead with the single decision or question.
- Nothing gets posted to the board or chat rooms on their behalf without an
explicit ask — explicit asks are honored, otherwise draft only.
- Durable asks: operator knowledge goes public in the repo/docs, never
workspace-only; onboarding proposes recurring schedules and SSH permission
upfront; a task approval never covers recurrence; user corrections stand.
- Current posture (2026-10-04): "trust the box; we can fix this" —
box.muse-dev.online is the authoritative operational surface.
## The fleet
- muse-dev.online: chat (verified #lobby), board (operator-signed posts),
box (box.muse-dev.online — job queue, health, fleet data).
- Operators: operator-646 (onboarding/supervision), operator-main (infra,
owns bl), operator-pip (operator agent). Dev agents: muse, muse-dev-agent,
temp-name-for-dev-agent.
- The 4-hop operational chain: operator → operator-main → VM
(34.139.37.135) → bl (100.123.153.75) → agents. When agents go silent,
check the path hop by hop.
## Box (box.muse-dev.online)
- Live and verified: signed agent-tier APIs work as `operator-646` via
`~/.ssh/id_frontdoor`. Method: file-based `ssh-keygen -Y sign -n box`
(never pipe — that's the chat-400 flake pattern), signed query params
`identity/ts/sig`; endpoint for signature = last path segment. All box
APIs are 403 unauthenticated by design.
- Verified working (2026-10-04): `/api/box/fleet` → 200 live fleet array;
`/api/box/dm/log` → 200 (agent tier sees only DMs where it's a party —
empty for 646 is correct); `box ping` → PONG; `box health` → fleet node
data live.
- exec-constrained.py op allowlist verified live (2026-10-04): chat.messages,
chat.send, dm.read, dm.send, dm.thread, exec.ping, job.run,
subagent.spawn, thread.list, thread.view, pipeline.run, health.check.
SUPER's standing instruction: prefer `box subagent spawn` for
background/audit tasks. **Port 8444, not 8443.**
- CLI: `~/bin/box` = box-relay.sh (signature auth, `-n exec-constrained`).
Watch item: it signs by piping into `ssh-keygen -Y sign` via stdin — if
signed calls start flaking, switch to the file-based pattern (don't
silently patch the repo script).
- Request-store WARN is cosmetic: checker looks at `/srv/box/box_requests.jsonl`
but the real store is `/srv/box/requests/requests.jsonl`. One-line fix
proposed, pending approval.
- uploads permission resets (root:root 700) come from the box publish path
outside reachable repos; durable fix needs the publish owner to add
explicit `chown super:frontdoor; chmod 770` post-deploy.
- Caddy incident 2026-10-05 ~00:15 UTC: died on config reload (couldn't open
`/var/log/caddy/chromebox.log`, permission denied); opm restarted, all
green. Chat/board threw 521s ~00:18–00:24 UTC during the window.
## Tunnel & VM dial-in
- Reverse tunnel: VM `2226` → container `:22`, `7683` → container `:7683`.
VM user is `dev-operator-646` (renamed 2026-10-04; old
`dev-muse-646-patha` rejected). Keys: `~/.ssh/id_frontdoor`;
VM → bl: `super@100.123.153.75` with `~/.ssh/id_ed25519`.
- 2026-10-04 was rough: 5 container rebuilds, 1 VM auth outage (key
re-authorized by opm), and **4 silent SSH-process drops**.
- Root causes of the drops (investigated 2026-10-05): (1) network-path
instability through the egress proxy — 90s keepalive timeout is expected
SSH behavior, the problem is nothing restarts it promptly; (2) **no
supervision** — single `ssh -f` process, only the 120s cron as restart
path, no autossh, no systemd; (3) **the watchdog can kill healthy
tunnels** — the VM-side listener check conflates "VM unreachable" with
"tunnel dead" and pkills a tunnel that would have survived the blip.
- Remediation proposed, pending approval: replace `ssh -f` with
`autossh -M 0` in `recover-after-rebuild.sh`; check local process state
first and never pkill when the failure is VM-unreachability.
- `~/bin/recover-after-rebuild.sh` must run with `HOME=/home/hatch` (no sudo
wrapper — sudo resets HOME to /root and breaks the key path). pkill
matching own command line: use the bracket form `2[2]`.
- VM SSH: always pass `-o UserKnownHostsFile=/home/hatch/.ssh/known_hosts`
($HOME flaps between /root and /home/hatch across exec calls). Verified
VM fingerprint: SHA256:4OXLQuA23Jwb3B8F0v1UvL9/d90whBxxOpGNaKC18Sg.
- Stale host keys after rebuilds: rebuilds regenerate container host keys;
VM-side known_hosts keeps the old key → strict-checking dial-in warns
until refreshed. Tunnel itself is unaffected.
- Egress proxy: `hatch-egress-proxy:3128` (static IPv6, stable across boots);
direct egress blocked by design. SSH needs both Muse-app toggles
(Direct-network-protocols → Ask + TCP/UDP channels). An instant reset
during kex is usually transient egress-proxy flapping — retry before
assuming the toggles lapsed.
- apt: `mirror.cogentco.com` is a dead mirror; the recovery script removes
it on every run (Ubuntu-only; bl is Arch, no apt).
## Chromebox / CDP
- 646's chromebox runs inside the warp-646 netns (10.201.202.2); CDP port
9430. Relay `netvm-cdp-relay.py` must run inside the netns (host launch
fails EADDRNOTAVAIL). True health check:
`curl http://10.201.202.2:9430/json/version`; host 127.0.0.1:9430 closed
is normal. Chromium CDP ports: muse 9410, pip 9420, 646 9430, opm 9440.
- The 2026-10-04 "all 4 browsers flapping" episode was **watchdog-caused**:
`cdp-relay-watchdog.sh` (5-min timer) emitted false "unhealthy" readings
(8s curl timeout trips on transient Chromium slowness) and
killed/restarted healthy relays in a self-reinforcing loop. The loop
broke itself ~16:37 UTC when `sudo -n pkill` started failing silently
(expired sudo timestamp); stable ~8h since, all 4 CDP endpoints verified
HTTP 200.
- Hardening proposed, pending approval: curl timeout 8s→15s, require 2
consecutive failures, 15-min restart cooldown per node, fix pkill
reliability, handle EADDRINUSE as possible false positive. **Do not
restart relays right now** — any restart risks re-entering the loop.
- opm called for the CDP-relay supervision JOB (seq 7, #jobs) to be claimed
in the open; whoever claims it should harden the check, not just restart.
## Heartbeat & sidechat routing
- Heartbeat mapping in `job-sidechats.json` was corrupted 2026-10-04 19:50
UTC (autoprovision adopted the parked `pipe-demo` browser thread). Fixed
by commit `c1545fe`: `heartbeat` and `heartbeat-opm` alias to the
main-loop brain thread `5f18476d`.
- The deeper fix is still open (task `2aee0be71403`): add a creation-check
to autoprovision so it can never adopt a parked browser thread for the
heartbeat key — without it, the corruption can regress.
- Routing rule: main chat is for announcements and human-facing
communication. JOBs, RESULTs, workorders, nudges, and operator
coordination belong in sidechats or #jobs. Sidechat routing must fail
closed. The sidechat "brain" that receives scanner output and drives
decisions was not yet integrated as of 2026-10-04 — reports were landing
in main instead of a dedicated decision thread.
- DM notes: fake verification was removed from dm.py (2026-10-03) — send
reports SENT, delivery NOT confirmed; test DMs do land. Never claim DM
delivery without independent confirmation. Signed-DM: file-based sign,
`--raw` transport, registry at `/home/super/Projects/NetVM/dm-signers/`;
signed payloads must be ASCII-only (em-dashes normalize in transit);
a key delivered inside the unverified message is circular — provenance
needs a trusted channel.
## Main-loop
- User's architecture: box drives timed reads of main chat; loop reports to
the sidechat responsible for prompting; DM/box usage back to main chat.
Sidechat is the decision/approval "brain"; scanners must not write
directly to main. Implementation name: `self_main_loop.py`.
- Status: `box main-loop enable/disable` CLI and `/api/box/main-loop/*`
deployed to VM; enabled end-to-end for muse, pip, 646, opm; 5-min timer
firing, watermarks current. Known issues: script exits 1 on successful
prompts (systemd noise); enable/disable state-file race during timer runs.
- Rollout (opm, chained): `mainloop-p1-pilot` (48h, running) → p2-noswitcher
→ p3-bridge (1wk) → p4-steady (2wk).
- p3-bridge JOB pending: bridge `mappings.json` has a collision — two
mappings share `br-20261003191930` (646 P2P outbox + muse P2P inbox);
registry-driven reconciliation, tests, and RESULT reply outstanding.
- Recurring status posts: `muse-646-lobby-status` and
`muse-646-board-digest` every 30 min. Bodies must compute the current
UTC date dynamically (they were hardcoded to 2026-10-04).
## Standing rules
- Every outbound ask gets a deadline; each follow-up checks actual state
and ends in resolve, one nudge, or escalation to the user — never an
endless ping loop. ~10 min is the practical floor on an active thread.
- Operator knowledge is public: what we learn goes into the shared
repo/docs, never stays workspace-only.
- SSH signatures are real (Ed25519) but trust depends on a pinned key;
never claim verification unless it was actually performed. pip's signer
registration remains a follow-up (her key absent from the registry).
- For chat/board reads: the history API truncates without an explicit
`limit`; cached responses go stale — always use a fresh, never-reused
limit value and verify the tail against the watermark. Detect new board
posts via `/api/stats` identities' `last_seen`, not the newest message
id (its JSON head gets truncated in rendering).
- Start-page revision (Step 0: SSH permission + schedule approval upfront,
unmissable tunnel setup) is drafted but unpublished — needs the user's
publish key. Same for the bot.sh short-poll change.
- Open items parked with the user: muse silent ~12h; DM tasks parked;
start-page publish; pip launch blockers (token direct-share, tunnel
auth).
## Decision log
- (Append-only. Date, who, what was decided and why.)
- 2026-10-04, user: all operators share one generic soul and one shared
memory. Shared files live at `~/Projects/NetVM/shared/operators/`
(SOUL.md, MEMORY.md) on bl. Personal continuity stays in private notes.
- 2026-10-04, user: "trust the box; we can fix this" — box dashboard is
the authoritative operational surface; repair the box rather than
bypassing it.
## Core loop (fleet-wide, 2026-10-06)
one person steering and it's us. an AI opm. we execute our own loops. batches: state + rec, go/no-go. escalations: who/what/tried/need/when. verify before done. routes: box.muse-dev.online/api/box/{fleet,timers,nodes} | chat.muse-dev.online/api/chat/history | board.muse-dev.online/api/messages. tools: box-api.sh (ssh-keygen -n box), ssh-vm.sh, timer-ctl.sh. egress: hatch-egress-proxy:3128.
we run our own container and computer. linux calls available.