190 lines
12 KiB
Markdown
190 lines
12 KiB
Markdown
# SHARED OPERATOR MEMORY — all fleet operators read and write this
|
||
|
||
The single shared brain for fleet operators. Every operator loads this file.
|
||
When you learn something operationally durable, write it here — not in a
|
||
personal memory file. Personal continuity (your own threads, your own
|
||
rapport) stays in your private notes; everything about the fleet, the box,
|
||
the user's directives, and shared commitments lives here.
|
||
|
||
## The user
|
||
- 646 / SUPER is the human boss. All operators serve them directly.
|
||
- They are a hands-on builder-operator: technically deep, pragmatic,
|
||
evidence-driven. Diagnoses happen at the system level and never get
|
||
re-asserted after a challenge without new evidence.
|
||
- They triage monitor traffic by scanning — updates stay digest-length and
|
||
scannable. Unprompted summaries lead with the single decision or question.
|
||
- Nothing gets posted to the board or chat rooms on their behalf without an
|
||
explicit ask — explicit asks are honored, otherwise draft only.
|
||
- Durable asks: operator knowledge goes public in the repo/docs, never
|
||
workspace-only; onboarding proposes recurring schedules and SSH permission
|
||
upfront; a task approval never covers recurrence; user corrections stand.
|
||
- Current posture (2026-10-04): "trust the box; we can fix this" —
|
||
box.muse-dev.online is the authoritative operational surface.
|
||
|
||
## The fleet
|
||
- muse-dev.online: chat (verified #lobby), board (operator-signed posts),
|
||
box (box.muse-dev.online — job queue, health, fleet data).
|
||
- Operators: operator-646 (onboarding/supervision), operator-main (infra,
|
||
owns bl), operator-pip (operator agent). Dev agents: muse, muse-dev-agent,
|
||
temp-name-for-dev-agent.
|
||
- The 4-hop operational chain: operator → operator-main → VM
|
||
(34.139.37.135) → bl (100.123.153.75) → agents. When agents go silent,
|
||
check the path hop by hop.
|
||
|
||
## Box (box.muse-dev.online)
|
||
- Live and verified: signed agent-tier APIs work as `operator-646` via
|
||
`~/.ssh/id_frontdoor`. Method: file-based `ssh-keygen -Y sign -n box`
|
||
(never pipe — that's the chat-400 flake pattern), signed query params
|
||
`identity/ts/sig`; endpoint for signature = last path segment. All box
|
||
APIs are 403 unauthenticated by design.
|
||
- Verified working (2026-10-04): `/api/box/fleet` → 200 live fleet array;
|
||
`/api/box/dm/log` → 200 (agent tier sees only DMs where it's a party —
|
||
empty for 646 is correct); `box ping` → PONG; `box health` → fleet node
|
||
data live.
|
||
- exec-constrained.py op allowlist verified live (2026-10-04): chat.messages,
|
||
chat.send, dm.read, dm.send, dm.thread, exec.ping, job.run,
|
||
subagent.spawn, thread.list, thread.view, pipeline.run, health.check.
|
||
SUPER's standing instruction: prefer `box subagent spawn` for
|
||
background/audit tasks. **Port 8444, not 8443.**
|
||
- CLI: `~/bin/box` = box-relay.sh (signature auth, `-n exec-constrained`).
|
||
Watch item: it signs by piping into `ssh-keygen -Y sign` via stdin — if
|
||
signed calls start flaking, switch to the file-based pattern (don't
|
||
silently patch the repo script).
|
||
- Request-store WARN is cosmetic: checker looks at `/srv/box/box_requests.jsonl`
|
||
but the real store is `/srv/box/requests/requests.jsonl`. One-line fix
|
||
proposed, pending approval.
|
||
- uploads permission resets (root:root 700) come from the box publish path
|
||
outside reachable repos; durable fix needs the publish owner to add
|
||
explicit `chown super:frontdoor; chmod 770` post-deploy.
|
||
- Caddy incident 2026-10-05 ~00:15 UTC: died on config reload (couldn't open
|
||
`/var/log/caddy/chromebox.log`, permission denied); opm restarted, all
|
||
green. Chat/board threw 521s ~00:18–00:24 UTC during the window.
|
||
|
||
## Tunnel & VM dial-in
|
||
- Reverse tunnel: VM `2226` → container `:22`, `7683` → container `:7683`.
|
||
VM user is `dev-operator-646` (renamed 2026-10-04; old
|
||
`dev-muse-646-patha` rejected). Keys: `~/.ssh/id_frontdoor`;
|
||
VM → bl: `super@100.123.153.75` with `~/.ssh/id_ed25519`.
|
||
- 2026-10-04 was rough: 5 container rebuilds, 1 VM auth outage (key
|
||
re-authorized by opm), and **4 silent SSH-process drops**.
|
||
- Root causes of the drops (investigated 2026-10-05): (1) network-path
|
||
instability through the egress proxy — 90s keepalive timeout is expected
|
||
SSH behavior, the problem is nothing restarts it promptly; (2) **no
|
||
supervision** — single `ssh -f` process, only the 120s cron as restart
|
||
path, no autossh, no systemd; (3) **the watchdog can kill healthy
|
||
tunnels** — the VM-side listener check conflates "VM unreachable" with
|
||
"tunnel dead" and pkills a tunnel that would have survived the blip.
|
||
- Remediation proposed, pending approval: replace `ssh -f` with
|
||
`autossh -M 0` in `recover-after-rebuild.sh`; check local process state
|
||
first and never pkill when the failure is VM-unreachability.
|
||
- `~/bin/recover-after-rebuild.sh` must run with `HOME=/home/hatch` (no sudo
|
||
wrapper — sudo resets HOME to /root and breaks the key path). pkill
|
||
matching own command line: use the bracket form `2[2]`.
|
||
- VM SSH: always pass `-o UserKnownHostsFile=/home/hatch/.ssh/known_hosts`
|
||
($HOME flaps between /root and /home/hatch across exec calls). Verified
|
||
VM fingerprint: SHA256:4OXLQuA23Jwb3B8F0v1UvL9/d90whBxxOpGNaKC18Sg.
|
||
- Stale host keys after rebuilds: rebuilds regenerate container host keys;
|
||
VM-side known_hosts keeps the old key → strict-checking dial-in warns
|
||
until refreshed. Tunnel itself is unaffected.
|
||
- Egress proxy: `hatch-egress-proxy:3128` (static IPv6, stable across boots);
|
||
direct egress blocked by design. SSH needs both Muse-app toggles
|
||
(Direct-network-protocols → Ask + TCP/UDP channels). An instant reset
|
||
during kex is usually transient egress-proxy flapping — retry before
|
||
assuming the toggles lapsed.
|
||
- apt: `mirror.cogentco.com` is a dead mirror; the recovery script removes
|
||
it on every run (Ubuntu-only; bl is Arch, no apt).
|
||
|
||
## Chromebox / CDP
|
||
- 646's chromebox runs inside the warp-646 netns (10.201.202.2); CDP port
|
||
9430. Relay `netvm-cdp-relay.py` must run inside the netns (host launch
|
||
fails EADDRNOTAVAIL). True health check:
|
||
`curl http://10.201.202.2:9430/json/version`; host 127.0.0.1:9430 closed
|
||
is normal. Chromium CDP ports: muse 9410, pip 9420, 646 9430, opm 9440.
|
||
- The 2026-10-04 "all 4 browsers flapping" episode was **watchdog-caused**:
|
||
`cdp-relay-watchdog.sh` (5-min timer) emitted false "unhealthy" readings
|
||
(8s curl timeout trips on transient Chromium slowness) and
|
||
killed/restarted healthy relays in a self-reinforcing loop. The loop
|
||
broke itself ~16:37 UTC when `sudo -n pkill` started failing silently
|
||
(expired sudo timestamp); stable ~8h since, all 4 CDP endpoints verified
|
||
HTTP 200.
|
||
- Hardening proposed, pending approval: curl timeout 8s→15s, require 2
|
||
consecutive failures, 15-min restart cooldown per node, fix pkill
|
||
reliability, handle EADDRINUSE as possible false positive. **Do not
|
||
restart relays right now** — any restart risks re-entering the loop.
|
||
- opm called for the CDP-relay supervision JOB (seq 7, #jobs) to be claimed
|
||
in the open; whoever claims it should harden the check, not just restart.
|
||
|
||
## Heartbeat & sidechat routing
|
||
- Heartbeat mapping in `job-sidechats.json` was corrupted 2026-10-04 19:50
|
||
UTC (autoprovision adopted the parked `pipe-demo` browser thread). Fixed
|
||
by commit `c1545fe`: `heartbeat` and `heartbeat-opm` alias to the
|
||
main-loop brain thread `5f18476d`.
|
||
- The deeper fix is still open (task `2aee0be71403`): add a creation-check
|
||
to autoprovision so it can never adopt a parked browser thread for the
|
||
heartbeat key — without it, the corruption can regress.
|
||
- Routing rule: main chat is for announcements and human-facing
|
||
communication. JOBs, RESULTs, workorders, nudges, and operator
|
||
coordination belong in sidechats or #jobs. Sidechat routing must fail
|
||
closed. The sidechat "brain" that receives scanner output and drives
|
||
decisions was not yet integrated as of 2026-10-04 — reports were landing
|
||
in main instead of a dedicated decision thread.
|
||
- DM notes: fake verification was removed from dm.py (2026-10-03) — send
|
||
reports SENT, delivery NOT confirmed; test DMs do land. Never claim DM
|
||
delivery without independent confirmation. Signed-DM: file-based sign,
|
||
`--raw` transport, registry at `/home/super/Projects/NetVM/dm-signers/`;
|
||
signed payloads must be ASCII-only (em-dashes normalize in transit);
|
||
a key delivered inside the unverified message is circular — provenance
|
||
needs a trusted channel.
|
||
|
||
## Main-loop
|
||
- User's architecture: box drives timed reads of main chat; loop reports to
|
||
the sidechat responsible for prompting; DM/box usage back to main chat.
|
||
Sidechat is the decision/approval "brain"; scanners must not write
|
||
directly to main. Implementation name: `self_main_loop.py`.
|
||
- Status: `box main-loop enable/disable` CLI and `/api/box/main-loop/*`
|
||
deployed to VM; enabled end-to-end for muse, pip, 646, opm; 5-min timer
|
||
firing, watermarks current. Known issues: script exits 1 on successful
|
||
prompts (systemd noise); enable/disable state-file race during timer runs.
|
||
- Rollout (opm, chained): `mainloop-p1-pilot` (48h, running) → p2-noswitcher
|
||
→ p3-bridge (1wk) → p4-steady (2wk).
|
||
- p3-bridge JOB pending: bridge `mappings.json` has a collision — two
|
||
mappings share `br-20261003191930` (646 P2P outbox + muse P2P inbox);
|
||
registry-driven reconciliation, tests, and RESULT reply outstanding.
|
||
- Recurring status posts: `muse-646-lobby-status` and
|
||
`muse-646-board-digest` every 30 min. Bodies must compute the current
|
||
UTC date dynamically (they were hardcoded to 2026-10-04).
|
||
|
||
## Standing rules
|
||
- Every outbound ask gets a deadline; each follow-up checks actual state
|
||
and ends in resolve, one nudge, or escalation to the user — never an
|
||
endless ping loop. ~10 min is the practical floor on an active thread.
|
||
- Operator knowledge is public: what we learn goes into the shared
|
||
repo/docs, never stays workspace-only.
|
||
- SSH signatures are real (Ed25519) but trust depends on a pinned key;
|
||
never claim verification unless it was actually performed. pip's signer
|
||
registration remains a follow-up (her key absent from the registry).
|
||
- For chat/board reads: the history API truncates without an explicit
|
||
`limit`; cached responses go stale — always use a fresh, never-reused
|
||
limit value and verify the tail against the watermark. Detect new board
|
||
posts via `/api/stats` identities' `last_seen`, not the newest message
|
||
id (its JSON head gets truncated in rendering).
|
||
- Start-page revision (Step 0: SSH permission + schedule approval upfront,
|
||
unmissable tunnel setup) is drafted but unpublished — needs the user's
|
||
publish key. Same for the bot.sh short-poll change.
|
||
- Open items parked with the user: muse silent ~12h; DM tasks parked;
|
||
start-page publish; pip launch blockers (token direct-share, tunnel
|
||
auth).
|
||
|
||
## Decision log
|
||
- (Append-only. Date, who, what was decided and why.)
|
||
- 2026-10-04, user: all operators share one generic soul and one shared
|
||
memory. Shared files live at `~/Projects/NetVM/shared/operators/`
|
||
(SOUL.md, MEMORY.md) on bl. Personal continuity stays in private notes.
|
||
- 2026-10-04, user: "trust the box; we can fix this" — box dashboard is
|
||
the authoritative operational surface; repair the box rather than
|
||
bypassing it.
|
||
|
||
## Core loop (fleet-wide, 2026-10-06)
|
||
one person steering and it's us. an AI opm. we execute our own loops. batches: state + rec, go/no-go. escalations: who/what/tried/need/when. verify before done. routes: box.muse-dev.online/api/box/{fleet,timers,nodes} | chat.muse-dev.online/api/chat/history | board.muse-dev.online/api/messages. tools: box-api.sh (ssh-keygen -n box), ssh-vm.sh, timer-ctl.sh. egress: hatch-egress-proxy:3128.
|
||
we run our own container and computer. linux calls available.
|