feat(hybrid-gateway): integrate muse-cli with Cloudflare netns isolation, symmetric sidechat routing, and 646-pip sync unblock

This commit is contained in:
operator
2026-10-04 22:54:01 +00:00
parent 3d5fbe5aeb
commit 1a271b1bbd
60 changed files with 5068 additions and 55 deletions
+77
View File
@@ -0,0 +1,77 @@
# DOM Headless Approvals Spec (muse.ai automation)
## Overview
When headless automation (`muse-chat-api.py` via CDP) drives a muse.ai
session, the browser may surface permission/confirmation dialogs that block
the flow (e.g. "Allow pip to share information with 34.139.37.135?").
This spec defines how the automation detects, classifies, and handles those
dialogs. Principle (from INFRA.md): **the chat IS the approval interface** —
no file-based queue; the automation signals when stuck and the operator
resolves it in conversation.
## Definitions
- **Approval dialog**: any in-DOM permission/confirmation prompt that gates
the automation's next action.
- **Trusted origin**: an IP in `TRUSTED_IPS` — our own infrastructure, where
auto-approval is safe. Current set:
- `34.139.37.135` — VM (gateway)
- `100.123.153.75` — bl (main compute)
- `100.81.31.9` — VM tailnet
- **APPROVAL_NEEDED**: the escalation signal. Printed to stderr as
`APPROVAL_NEEDED: <dialog summary>`, process exits with code **2**.
## Detection
`check_approvals(ws)` evaluates in the page DOM:
1. Body text containing both `Allow` and `to share` → permission prompt.
Narrows to elements whose innerText contains both and is < 500 chars.
2. Two or more buttons whose text includes `allow`, `deny`, or `block`
→ likely permission dialog; captures the closest container's text.
Returns a list of `(dialog_text, is_trusted, action_taken)`.
## Classification
Extract IPv4 addresses from the dialog text. The dialog is **trusted** iff
any extracted IP is in `TRUSTED_IPS`. Dialogs with no recognizable IP are
**untrusted** (fail closed).
## Handling
- **Trusted**: auto-approve by clicking the button whose text contains
`allow once`, else the button whose text is exactly `allow`.
Records `clicked:<button text>`.
- **Untrusted**: do NOT click. Return `APPROVAL_NEEDED`; the calling command
prints the summary to stderr and exits 2.
## Enforcement points
| Command | Behavior |
|---|---|
| `send` | Checks approvals first. Untrusted dialog → `APPROVAL_NEEDED`, exit 2, message NOT sent. |
| `messages` | Non-blocking check (observes, does not gate). |
| `wait` | Re-checks every 5s during the wait. Untrusted dialog → `APPROVAL_NEEDED`, exit 2. |
| `approvals` | Reports pending dialogs with trusted/action status. Never gates. |
## Exit-code contract
| Code | Meaning |
|---|---|
| `0` | Done. |
| `1` | Failed (bad args, no page, send error, …). |
| `2` | `APPROVAL_NEEDED` — human input required; safe to retry after resolution. |
## Operator flow
1. Automation exits 2 with `APPROVAL_NEEDED: <details>`.
2. Operator relays the details to the user via chat.
3. User provides the needed input (e.g. confirms, provides OTP).
4. Operator re-runs the command; the dialog is either gone or now trusted.
## Non-goals (explicitly out of scope)
- **DM content provenance**: this spec covers *browser dialogs encountered by
automation*. It does NOT authenticate who authored a DM. A message injected
via `dm.py send` carries no signature; a receiving agent cannot verify the
claimed sender from the tooling. (Observed 2026-10-03: pip's agent rightly
refused to act on an unsigned DM.) Provenance is a separate spec.
- **Credential approval**: Meta account write operations use the human-gated
flow in `docs/META-ACCOUNTS-API.md`, not this spec.
## Future work
- Allowlist dialog *types* (not just IPs) for finer auto-approval.
- Structured `APPROVAL_NEEDED` payloads (JSON) for machine-readable relay.
- DONE 2026-10-03: dm-level provenance via ssh-keygen signing. `dm-sign.sh` signs with `ssh-keygen -Y sign -n dm` (file-based); `dm.py send --raw` transports the signed block verbatim; `dm.py verify-sig` verifies against `dm-signers/<sender>.pub` via `ssh-keygen -Y verify`.
+94
View File
@@ -0,0 +1,94 @@
# Hybrid Headless Gateway & Chromebox Architecture: `muse-cli` Adaptation
## Overview
NetVM previously automated agent interaction (thread reading, message sends, job execution) solely via headless Chromebox containers over Chrome DevTools Protocol (CDP) and DOM manipulation.
While browser automation is indispensable for interactive UI flows (initial authentication, OTP submission, age verification, visual inspection), routine agent messaging and thread polling over DOM mutation suffered from high latency and UI placement flakiness (`placement_failed`, `NOCOMPOSE`, `NOINPUT`).
To resolve this, we adapted **`muse-cli`** into the NetVM `box` (`super-cli.py`) ecosystem as a **fast, headless gateway transport** connecting directly over WebSockets (`wss://gateway.muse.ai/v1/noise`) via encrypted Noise protocol frames (`Noise_XX_25519_AESGCM_SHA256`).
---
## 1. Strict Cloudflare Identity & Egress Isolation
A core architectural invariant of NetVM is that **no fleet node shares an unisolated egress network identity or ambient host IP**. Every command executed via the CLI maintains strict network and session separation.
### Network Namespace Boundary
Each node possesses a dedicated Linux network namespace (`warp-<node>`) containing its own WireGuard interface (`wb-<tag>`), routing all traffic through its assigned Cloudflare WARP client identity:
- `warp-muse` (Interface: `wb-4016c3db`)
- `warp-pip` (Interface: `wb-fed5038b`)
- `warp-646` (Interface: `wb-ed0b853b`)
- `warp-opm` (Interface: `wb-54353f9c`)
### Execution Wrapper: `bin/muse-cli-node`
The executable wrapper [`bin/muse-cli-node`](file:///home/super/Projects/NetVM/bin/muse-cli-node) bridges CLI requests into the agent's isolated namespace:
```bash
bin/muse-cli-node <node> <subcommand> [args...]
```
1. **Network Egress**: Dispatches execution through [`bin/netvm-exec.sh`](file:///home/super/Projects/NetVM/bin/netvm-exec.sh), ensuring all WebSocket and HTTP requests originate from the node's specific WARP interface.
2. **Session Cookie Partitioning**: Uses isolated cookie stores located in `~/.config/muse-cli/<node>/cookies.txt` (permissions `0600`).
3. **Auto-Healing Cookie Refresh**: On receiving an `AuthError` (or 401 Unauthorized), `muse-cli-node` intercepts the error, runs [`bin/refresh-node-cookies.py`](file:///home/super/Projects/NetVM/bin/refresh-node-cookies.py) to export fresh session cookies directly from the node's local Chromium instance via CDP inside its netns, and transparently retries the command once.
---
## 2. Hybrid Modality Architecture
NetVM leverages both modalities according to their operational strengths:
| Task / Domain | Primary Modality | Fallback / Complementary Modality |
| :--- | :--- | :--- |
| **Routine Thread Listing** | Gateway (`muse_hybrid.get_threads`) | DOM CDP (`box-chat.py thread-list`) |
| **Thread History Reading** | Gateway (`muse_hybrid.get_history`) | DOM CDP (`box-chat.py thread-messages`) |
| **High-Frequency Job/DM Dispatch** | Gateway (`muse-cli send`) | DOM CDP (`muse-chat-api.py send`) |
| **Live Event Streaming** | Gateway (`muse-cli watch`) | Siphon log polling (`siphon-bl.py`) |
| **Initial Login & OTP Entry** | DOM CDP ([`bin/muse-signin.py`](file:///home/super/Projects/NetVM/bin/muse-signin.py)) | N/A (Requires interactive DOM inputs) |
| **Visual Screencast & Audit** | DOM CDP (`Page.startScreencast`) | N/A |
---
## 3. Programmatic & CLI Interfaces
### Python Interface: `bin/muse_hybrid.py`
Provides clean programmatic access for scripts and background daemons:
```python
import muse_hybrid
# Fetch thread listing (sub-second)
threads, err = muse_hybrid.get_threads("pip")
# Fetch message history
msgs, err = muse_hybrid.get_history("pip", thread_id="4466d0c1-7961-4cf3-b99d-1ab7c38484c2", limit=5)
# Fast send
res, err = muse_hybrid.send_message("pip", "Hello from gateway", thread_id="4466d0c1-...")
```
### CLI Interface: `super-cli.py` (`box`)
1. **Direct Gateway Passthrough**:
Operators can execute any `muse-cli` command under a specific node's identity:
```bash
box muse pip status
box muse 646 threads
box muse opm watch
```
2. **Accelerated Thread Commands**:
`box thread list <agent>` and `box thread view <agent> <thread_id>` automatically query `muse_hybrid` first. If the gateway encounters an issue, they gracefully fall back to the existing `box-chat.py` CDP path.
---
## 4. Verification & Health Audit
To verify the hybrid gateway across all fleet nodes:
```bash
# Verify status per node
for node in muse pip 646 opm; do
echo "=== $node ==="
box muse $node status | jq -r '.identity.name'
done
# Verify thread view via hybrid path
box thread view pip 4466d0c1-7961-4cf3-b99d-1ab7c38484c2 --limit 3
```
+78
View File
@@ -0,0 +1,78 @@
# Identity Variance Testing
## Why
All NetVM nodes currently egress from a single Cloudflare IP (`104.28.195.181`),
but each node holds a *distinct* WireGuard identity. When production rate limits
hit, we need data to answer: per-identity, per-IP, time-based, or random?
This framework builds that dataset gently: ~2 probes per node every 10 minutes.
## Files
- `bin/identity-variance-test.py` — probe harness + analyzer
- `.state/identity-variance/variance.jsonl` — results log (git-ignored, append-only)
- (proposed, not installed) systemd unit + timer for 10-min cadence
## What it measures
Per probe, per node:
- `http_status` — 200s vs 429 (rate limited) vs 403 (blocked) vs curl failures
- `response_ms` — full-request timing from inside the netns
- `egress_ip` — IP seen from inside the netns (per node, per cycle)
- `identity_hash` — sha256 (truncated) of the node's WireGuard private key.
Detects identity rotation; the key itself is NEVER logged.
- `remote_ip` — the IP curl actually connected to
Probes (both lightweight, no auth, no writes):
1. `https://www.cloudflare.com/cdn-cgi/trace` — what Cloudflare sees
2. `https://muse.ai/` — production-relevant landing page
Traffic budget: 5 nodes × 2 probes every 10 min ≈ 6 req/hr per node. Negligible.
## Log format (JSONL)
```json
{"ts": "2026-10-04T21:50:00+00:00", "node": "muse", "probe": "cf-trace",
"url": "https://www.cloudflare.com/cdn-cgi/trace", "curl_rc": 0,
"http_status": 200, "response_ms": 312.4, "bytes": 240,
"remote_ip": "104.16.0.0", "rate_limited": false, "blocked": false,
"ok": true, "elapsed_wall_ms": 415.2,
"identity_hash": "a43dded8edc358e2", "egress_ip": "104.28.195.181"}
```
## Usage
```bash
# one manual cycle
bin/identity-variance-test.py probe
# analyze (all data)
bin/identity-variance-test.py analyze
# last 24h only
bin/identity-variance-test.py analyze --since 24
```
## Reading the analysis
- `http_429` differing by node → per-identity differential treatment
- all nodes 429 at the same `ts` → correlated (per-IP) throttling
- 429s clustering by time of day → time-based limits
- uniform, rare 429s with no pattern → effectively random
## Installing the timer (human-gated)
```bash
sudo cp systemd/identity-variance.{service,timer} /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now identity-variance.timer
systemctl list-timers | grep identity-variance
```
## Notes
- `sudo -n ip netns exec` is required; script fails cleanly otherwise.
- The `def` node is included but will report `node-skip` until assigned.
- Log rotation: JSONL grows ~1.5KB/cycle (~200KB/day). Rotate monthly:
`mv variance.jsonl variance-$(date +%Y%m).jsonl`.
+53
View File
@@ -0,0 +1,53 @@
# Per-Node Rate Limit Policy (2026-10-04)
## Why
All 4 fleet nodes (muse, pip, 646, opm) egress from a single Cloudflare IP
(`104.28.195.181`). Cloudflare therefore sees the whole fleet as one client.
When fleet timers fire simultaneously, synchronized send bursts look like one
aggressive client and invite correlated rate-limiting — the likely shape of
the "all 4 browsers flapping" episodes.
Per-agent rate limiting with deterministic jitter de-correlates the fleet:
each node's send phase drifts apart instead of hitting in lockstep.
## Policy
| Agent | Base interval | Jitter factor | Effective interval | Burst | Max/min |
|---|---|---|---|---|---|
| muse | 3.0s | 1.2917 | 3.875s | 5 | 20 |
| pip | 3.0s | 0.8183 | 2.455s | 5 | 20 |
| 646 | 3.0s | 1.1193 | 3.358s | 5 | 20 |
| opm | 3.0s | 0.8287 | 2.486s | 5 | 20 |
Mean effective interval ≈ 3.04s — the jitter does **not** make limits
stricter on average; it only spreads phases.
- **Jitter:** ±30% deterministic per agent, seeded by
`sha256("netvm-rate-jitter:v1:" + agent)`. Same agent → same factor every
run (reproducible); different agents → different factors (de-correlated).
Uses an isolated `random.Random`, never touches global random state.
- **Scope:** per-agent token buckets. State lives in one shared file
(`/tmp/netvm-rate-limit.json`) but keys are namespaced by agent
(`{agent}`, `{agent}_recent`) — buckets are fully independent.
- **Enforcement points:**
- `job-dispatch.py` — `rate_limit_wait(agent)` / `rate_limit_wait(sender_agent)`
(pre-existing; now jittered, no call-site change needed)
- `dm.py::dm_send` — `rate_limit_wait(agent)` before send, keyed on the
sender (added 2026-10-04). Near no-op at normal pacing since sends
already sleep 2–3s; it only bites on burst loops.
## Inspect
```bash
python3 bin/rate_limiter.py --policy # effective policy per agent
python3 bin/rate_limiter.py --policy opm --no-jitter
```
## Notes
- `/tmp` state does not survive reboots; buckets start empty (permissive).
- To disable jitter for a caller: `rate_limit_wait(agent, jitter=0)`.
- Distinct egress IPs per node (Cloudflare Zero Trust dedicated egress)
remains the structural fix; this policy is the de-correlation layer that
works regardless.
+182
View File
@@ -0,0 +1,182 @@
# Thread Bookkeeping with Pin & Archive
**Status:** runbook (2026-10-04)
**Applies to:** muse.ai side-chat threads used by the fleet DM system
## Why bookkeeping matters
Side-chat threads are addressed by UUID. UUIDs go stale (deletion,
account-side cleanup, platform reissue — observed 2026-10-04 when
`heartbeat-opm` changed UUIDs in a single day and when the hardcoded
`1e75a740` "646 tasks" UUID died ~90 minutes after being committed).
Stale UUIDs cause silent misdelivery: the SPA redirects to an arbitrary
valid thread instead of 404ing, and the old placement-blind verify
claimed success.
Duplicates compound the problem. On 2026-10-04 the user observed two
`BRAIN-INIT` messages, indicating two brain threads where one was
expected. Fuzzy sidebar name search then matches the wrong thread.
Pin and archive are the native muse.ai mechanisms for managing this.
They are gateway API methods (`/api/session/pin`,
`/api/session/unpin`, `/api/session/archive`,
`/api/session/unarchive`, `/api/session/delete`,
`/api/session/rename`), **not** CDP browser actions. Our
`muse-chat-api.py` (CDP-based) does not implement them.
## Tooling: muse-cli
The supported path is the third-party `muse-cli` tool
(`nikships/muse-cli`), which speaks the gateway protocol (Noise XX +
protobuf) directly. It is **not** installed on bl, the VM, or the
operator container as of 2026-10-04.
### Install (on bl)
```bash
# Node.js required
npm install -g muse-cli # or per-repo install per muse-cli README
```
### Auth
muse-cli needs a `hatch_sess` cookie from a logged-in browser session.
Obtain it from the operator's own browser profile — never from another
agent's session. Treat the cookie as a credential: transient use, never
logged, never in chat.
### Commands (per muse-cli PROTOCOL.md)
```bash
muse-cli pin <session_id> # pin a thread
muse-cli unpin <session_id> # unpin a thread
muse-cli session-archive <session_id> # archive a thread
muse-cli session-unarchive <session_id> # restore an archived thread
muse-cli session-rename <session_id> <title> # rename a thread
muse-cli session-delete <session_id> # delete a thread (destructive)
```
`<session_id>` is the thread UUID (the hex string in
`https://muse.ai/thread/<uuid>`).
## When to pin a thread
Pin threads that are **critical infrastructure** or **frequently
accessed during triage**:
| Thread | Alias / reuse_key | Why pinned |
|---|---|---|
| main-loop brain | `main-loop brain` | Fleet coordination point |
| 646 tasks | `646 tasks` | 646's primary work thread |
| heartbeat | `heartbeat-opm` | Pipeline health signal |
| 646-pip-coord | `646-pip-coord` | Cross-operator coordination |
| muse tasks | `muse tasks` | muse's primary work thread |
Pinning keeps these at the top of the sidebar, which:
- Prevents fuzzy-name-search confusion (searching "646" is less likely
to match the wrong thread when the right one is pinned and visible).
- Makes important threads easy to find during incident triage.
- Does **not** change the UUID — pin is purely a display/ordering
operation. `job-sidechats.json` mappings are unaffected.
## When to archive a thread
Archive threads that are **stale, duplicate, test, or dead**:
- **Duplicates:** the double `BRAIN-INIT` incident — archive the
surplus thread, keep the registered one.
- **Stale UUIDs:** threads whose UUID no longer resolves (e.g. the dead
`1e75a740` if it still exists on 646's account).
- **Test threads:** `test-auto-prov` (`b96dd020`), `pipe-7d2896`,
`pipe-1b4579`, and other autoprovision experiments.
- **Superseded threads:** old threads replaced by a renamed or
re-provisioned canonical thread.
Prefer **archive** over **delete**: archiving is reversible
(`session-unarchive`), deleting is not. Delete only when certain the
thread holds nothing worth recovering.
## Listing pinned vs archived threads
muse-cli's session listing distinguishes state. Our CDP-based
`sidechat list` scrapes visible sidebar text only and **cannot**
distinguish pinned/archived/deleted threads — another reason to use
muse-cli for bookkeeping operations.
After any pin/archive operation, verify with the list command and
confirm the expected state before updating mappings.
## Integration with job-sidechats.json
`job-sidechats.json` (on bl) is the dynamic UUID registry:
`{ "<alias>": { "thread_uuid": ..., "agent": ..., ... } }`.
Rules:
1. **Pin changes nothing** in `job-sidechats.json`. The UUID is the
same; only sidebar ordering changes.
2. **Archive the registered thread → update the mapping.** Either
remove the entry (so the next send autoprovision a fresh thread) or
point it at the surviving canonical thread's UUID.
3. **Archive a non-registered duplicate → no mapping change needed**,
but note it in the commit message for traceability.
4. **Never hardcode UUIDs** in `SIDCHAT_ALIASES` or elsewhere (lesson
from 2026-10-04). Use name-based autoprovision with the dynamic
registry. A hardcoded UUID is a ticking time bomb: it *will* go
stale, and there is no alert when it does.
## Case study: duplicate brain INIT (2026-10-04)
**What happened.** The user observed two `BRAIN-INIT` messages in the
UI, indicating two brain threads where the design calls for one. The
brain thread (`5f18476d-8994-49e7-a9e0-4838732363fe`, created 19:40 UTC
on opm's account with placement confirmed) was never registered in
`job-sidechats.json`, so 646 reported it as "still pending" and
"unwired" while a second INIT appeared.
**Contributing factors.**
- The brain thread was created but the registration step was skipped.
- Without registration, later code paths could not find the canonical
thread and created (or appeared to create) another.
- No pin/archive pass had ever been run, so there was no canonical
"this is the one true brain" signal in the sidebar.
**How archive fixes it.**
1. Identify both thread UUIDs (from dm-log `alias_resolved` /
`nav_ok` records or muse-cli session list).
2. Confirm which UUID is registered in `job-sidechats.json` under
`"main-loop brain"` — that one is canonical.
3. `muse-cli session-archive <duplicate_uuid>` on the surplus thread.
4. Pin the canonical brain thread so it stays visible.
5. Commit the `job-sidechats.json` state with a note referencing the
archived duplicate.
**Prevention.** After creating any canonical thread: register it in
`job-sidechats.json` immediately, pin it, and note the UUID in the
commit message. Run a periodic (monthly) archive pass over test and
stale threads.
## Quick reference
```bash
# one-time: install + auth (hatch_sess cookie, transient)
# pin the critical set (opm account, adjust per agent)
muse-cli pin 5f18476d-8994-49e7-a9e0-4838732363fe # main-loop brain
# archive a duplicate / stale thread
muse-cli session-archive <duplicate_uuid>
# verify, then update job-sidechats.json on bl if the archived
# thread was the registered one (re-point or remove the entry)
```
## Open items (2026-10-04)
- muse-cli is not yet installed anywhere in the fleet.
- `hatch_sess` cookie handling needs a defined transient-use procedure.
- No scheduled archive pass exists yet; propose monthly.
- CDP-based `sidechat list` cannot see archived state — bookkeeping
must go through muse-cli, not dm.py.
+100
View File
@@ -0,0 +1,100 @@
# Timer Stagger — Fleet Polling De-correlation
**Date:** 2026-10-04 ~21:50 UTC
**Why:** All 4 NetVM nodes egress via the same Warp IP (104.28.195.181). Synchronized
timers = 4x correlated traffic bursts from one IP = correlated rate-limit risk.
Staggering spreads load; combined with per-node rate limiting it de-correlates
the fleet's egress signature.
## Before
Three heaviest muse.ai-egress timers fired at the **exact same second** every 5 min
(`OnCalendar=*:0/5`): `job-heartbeat`, `self-main-loop`, `main-chat-watchdog`.
`accounts-health` + `loop-remediator` coincided every 15 min. All four `wake-*`
timers fired hourly at :00. Three of four `chromebox-watchdog` timers were
synchronized to the second.
## After — stagger schedule
### bl user timers (`~/.config/systemd/user/`, hand-managed)
| Timer | Old | New | Fires at |
|---|---|---|---|
| job-heartbeat | `*:0/5` | `*:0/5` (phase 0, unchanged) | :00 :05 :10 … |
| self-main-loop | `*:0/5` | `*:1/5` | :01 :06 :11 … |
| main-chat-watchdog | `*:0/5` | `*:2/5` | :02 :07 :12 … |
| keepalive (box-managed) | `*:0/5` | `*:4/5` | :04 :09 :14 … |
| accounts-health | `*:0/15` | `*:3/15` | :03 :18 :33 :48 |
| loop-remediator | `*:0/15` | `*:8/15` | :08 :23 :38 :53 |
| followup-sweeper | `*:*` (@:00s) | `*:*:20` | :20s of every minute |
| agent-health | OnBootSec/OnUnitActiveSec 5m | + `RandomizedDelaySec=60s` | jittered ±60s |
| fleet-alert-check | OnBootSec/OnUnitActiveSec 5m | + `RandomizedDelaySec=60s` | jittered ±60s |
| exec-watch | OnUnitActiveSec 15m | + `RandomizedDelaySec=120s` | jittered ±120s |
systemd natively supports offset calendar specs (`*:1/5`); no wrapper scripts needed.
### Box-managed timers (`/srv/box/timers.json` → `timer-sync.py --all`)
| Timer | Old | New |
|---|---|---|
| wake-muse | `*-*-* *:00:00` | `*-*-* *:05:00` |
| wake-pip | `*-*-* *:00:00` | `*-*-* *:20:00` |
| wake-646 | `*:0/30` | `*-*-* *:35:00` |
| wake-opm | `*:0/30` | `*-*-* *:50:00` |
| loop-health | `*:0/15` | `*:11/15` |
| keepalive | `*:0/5` | `*:4/5` |
(`wake-646`/`wake-opm` were actually on `*:0/30`, not hourly — now spread 15 min apart
across the hour like the others.)
### System timers (`/etc/systemd/system/`, hand-managed)
| Timer | Change |
|---|---|
| chromebox-watchdog-{pip,646,muse,opm} | + `RandomizedDelaySec=30s` (local checks only; de-correlates restart bursts) |
## Backups (pre-change)
- bl user timers: `~/.config/systemd/user/bak-stagger-20261004/` (12 files)
- system timers: `/root/bak-stagger-20261004/`
- timers.json: `/srv/box/timers.json.bak-stagger-20261004`
## Verification (2026-10-04 ~21:52 UTC)
`systemctl --user list-timers` and `systemctl list-timers` confirm the new phases:
followup-sweeper @ :50:20, main-chat-watchdog @ :52:00, loop-remediator @ :53:00,
accounts-health @ 22:03:00, keepalive @ :54:00, wake-muse @ 22:05, wake-pip @ 22:20,
wake-646 @ 22:35, wake-opm @ 22:50. Watchdogs spread over ~7s via jitter.
Note: `self-main-loop` did one immediate catch-up run at 21:49:21 on reload
(`Persistent=true` treating the 21:46 slot as missed under the new phase) — benign,
one extra run.
## Rollback
```bash
# bl user timers
cp ~/.config/systemd/user/bak-stagger-20261004/*.timer ~/.config/systemd/user/
systemctl --user daemon-reload
# system timers
sudo cp /root/bak-stagger-20261004/*.timer /etc/systemd/system/
sudo systemctl daemon-reload
# box timers
cp /srv/box/timers.json.bak-stagger-20261004 /srv/box/timers.json
/srv/box/bin/timer-sync.py --all
```
## Not changed (deliberately)
- `box-http-health` (`*:0/15`) / `box-service-health` (`*:7/15`) — already staggered.
- `siphon`, `cdp-relay-watchdog`, `shadow`, `backup-sync` — unrelated cadences.
- `chat-state`, `response-harvester` — inactive.
- `self-main-loop` still touches all 4 browsers inside one run — staggering the
*timer* doesn't de-correlate *within* the run. Follow-up: add inter-node sleeps
inside `self_main_loop.py` itself.
## Status
Live on bl + VM as of 2026-10-04 ~21:52 UTC. Doc uncommitted (review-then-commit);
timer unit files live outside the NetVM repo (`~/.config/systemd/user`,
`/etc/systemd/system`) — consider versioning them if drift recurs.
+238
View File
@@ -0,0 +1,238 @@
# WARP Egress Fix — Distinct Egress IPs per NetVM Node
Date: 2026-10-04
Status: DESIGN (not yet implemented)
Author: operator-main
## 1. Goal
Give each NetVM node (`muse`, `pip`, `646`, `opm`, `def`) its own distinct
public egress IP, so that a Cloudflare rate-limit / block / partition on one
node's egress IP does not take down the whole fleet at once.
Background: the 2026-10-04 partitioning investigation found all nodes egress
from **one shared IP (104.28.195.181)**. Cloudflare sees the fleet as a single
client; a single throttle/block correlates into a fleet-wide outage — which
matches the observed "all 4 browsers flapping" pattern better than four
independent failures.
## 2. Empirical finding (corrects the premise)
**Each node ALREADY has a distinct Warp identity** — and still shares one
egress IP.
Verified live 2026-10-04 ~21:40 UTC on bl:
| node | identity (sha256 of PrivateKey, truncated) | egress IP now |
|------|--------------------------------------------|---------------|
| muse | 0fce50a5aff0 | 104.28.195.181 |
| pip | e9fbb39cc4c5 | 104.28.195.181 |
| 646 | e764d9d80fa8 | 104.28.195.181 |
| opm | 771725bf8a5b | 104.28.195.181 |
| def | fa868411c422 | 104.28.195.181 |
Identities live at `/etc/netvm/<node>.conf` (root-owned, 0600), one per node,
generated via `netvm-new-identity.sh`. Egress IPs probed per-node with
`ip netns exec warp-<node> curl -s https://api.ipify.org`.
**Why distinct identities share one egress IP:** consumer Cloudflare Warp
(wgcf-registered) exits through Cloudflare's anycast edge. The egress IP is
determined by the serving PoP/edge, not by the client identity — it is drawn
from a pool shared across many Warp users. Registering more identities does
not change the egress IP.
**Consequence:** "generate 4 distinct Warp identities" is already done and is
NOT the fix. The fix must change *how traffic exits*, not *who the client is*.
## 3. Options
### Option A — Cloudflare Zero Trust dedicated egress IPs (paid, native)
Cloudflare Zero Trust (WARP for Teams) offers **dedicated egress IPs** as a
paid add-on. With one dedicated egress IP per node and Gateway egress
policies matched on device identity, each node gets a stable, distinct,
organization-owned egress IP.
- Pros: stays on Cloudflare/WireGuard; `netvm-node-up.sh` flow mostly
unchanged (different endpoint + enrollment cert); stable IPs you control;
per-node egress policies give true isolation.
- Cons: paid plan + per-IP add-on (human decision + payment); replaces
wgcf consumer identities with Zero Trust enrollment — the largest config
change of the four options; enrollment certs become new credentials to
guard (same transient-handling rules as Warp keys).
- Human-gated: yes (purchase + org setup).
### Option B — Per-node commercial VPN providers (distinct IPs, non-Warp)
Give each node a WireGuard config from a *different* commercial VPN provider
(e.g. Mullvad, ProtonVPN, IVPN), each with its own exit IP.
- Pros: genuinely distinct egress IPs, guaranteed by construction; no
Cloudflare correlation at all.
- Cons: one account + credentials per provider (5 accounts); non-uniform
configs; recurring cost per provider; abandons the Warp uniformity the
fleet was built on; each provider's ToS / automation posture must be
vetted.
- Human-gated: yes (accounts + payment).
### Option C — Accept shared egress; de-correlate and detect (no new infra)
Keep consumer Warp. Fix the *failure mode* instead of the IP:
1. **Warp egress health stage in `chromebox-watchdog.sh`** (5th stage after
the existing 4): check `wg show <WG> latest-handshakes` age AND probe
egress from inside the netns
(`ip netns exec warp-<node> curl -m5 https://www.cloudflare.com/cdn-cgi/trace`).
On stale handshake or failed probe set
`HEALTH_FAIL_REASON="warp egress down"` and take the same relaunch path
as a dead CDP. This converts the invisible partition (CDP green, browser
offline) into a visible, recoverable event.
2. **React-readiness "match" assertion** in `muse-chat-api.py` before
send/verify: assert title contains `Chat —` and the message-list
container (`role="log"`) is present and non-skeleton. Fail closed
otherwise. (A partitioned browser otherwise passes naive readiness.)
3. **Per-node rate limiting + jitter** on muse.ai-bound traffic so the four
nodes don't look like one coordinated client. `bin/rate_limiter.py`
already exists — ensure token buckets are per-node, not shared, and add
jitter to the 5-minute polling loops.
4. **Stagger the fleet polling loops** (heartbeat, main-loop digests,
health checks) so four nodes never fire in lockstep.
- Pros: no cost, no new credentials, no account changes; directly fixes the
observed failure (invisible partition) and the most likely trigger
(correlated rate-limiting); fully reversible.
- Cons: egress IP stays shared — a hard Cloudflare block of 104.28.195.181
still hits all nodes (but now it is *detected*, *logged distinctly*, and
*recovered from* instead of flapping silently).
- Human-gated: no.
### Option D — Hybrid: per-node egress proxies for sensitive paths only
Keep Warp for general traffic; route muse.ai-bound traffic per node through
distinct SOCKS5/HTTP proxies with distinct IPs (browser `--proxy-server`
per profile).
- Pros: distinct IPs where it matters (muse.ai automation); rest of stack
untouched.
- Cons: proxy accounts + credentials per node; added latency on the hot
path; proxy reliability becomes a new failure domain; Chromium proxy
config per profile adds bookkeeping.
- Human-gated: partially (proxy accounts).
## 4. Recommendation
**Implement Option C now.** It is the only option that fixes the failure we
actually observed (partitions invisible to the watchdog) with zero new cost
or credentials, and it is fully reversible.
**Keep Option A as the paid path** if truly distinct, stable egress IPs are
later required (e.g. Cloudflare starts hard-blocking the shared egress IP,
or per-node IP reputation becomes a product need). Option A is the
architecturally clean answer; it just needs a human purchase decision.
**Do not pursue Option B** unless the fleet leaves Cloudflare for other
reasons — the ops overhead (5 VPN accounts) outweighs the benefit while
Option C mitigates the correlation.
## 5. Identity rotation procedure (for reference)
Not needed for this fix (identities are already distinct), but recorded
because the task asked how identities are generated. Per AGENTS.md, the user
has explicitly authorized operators to run `netvm-new-identity.sh`
("Warp credentials are NOT top secret").
One node at a time, rolling:
```bash
# on bl, as super (sudo -n via allowlist for netvm-node-up.sh)
NODE=muse # repeat per node
sudo cp /etc/netvm/${NODE}.conf /root/netvm-backup-$(date +%F)/${NODE}.conf
sudo rm /etc/netvm/${NODE}.conf
sudo /home/super/Projects/NetVM/bin/netvm-new-identity.sh ${NODE}
sudo /home/super/Projects/NetVM/bin/netvm-node-up.sh ${NODE}
ip netns exec warp-${NODE} curl -s -m 10 https://api.ipify.org; echo " <- ${NODE}"
```
`netvm-node-up.sh` is idempotent: it rebuilds the netns, veth pair,
WireGuard interface, and CDP relay from the conf. Wait for the new tunnel's
handshake (`wg show wb-<tag> latest-handshakes`) before moving to the next
node. Never rotate two nodes simultaneously — keep 4/5 of the fleet up.
## 6. Migration plan (Option C)
Order matters: instrument first, then de-correlate.
1. **Watchdog 5th stage** — edit `bin/chromebox-watchdog.sh` `healthy()`:
add Warp handshake-age check + in-netns egress probe; set
`HEALTH_FAIL_REASON="warp egress down"`. Deploy to one node (opm) first,
watch one full 5-minute cycle, then roll to all nodes. Commit.
2. **React-readiness assertion** — add the match check in `muse-chat-api.py`
before send/verify paths. Test against a known-good and a known-bad
(partitioned or landing-page) browser. Commit.
3. **Per-node rate limiter audit** — confirm `bin/rate_limiter.py` buckets
are keyed per node; add jitter to fleet polling loops (heartbeat timer,
self-main-loop.timer, health checks) so nodes don't fire in lockstep.
Commit.
4. **Stagger verification** — after deploy, confirm via logs that the four
nodes' 5-minute loops are spread across the minute, not clustered.
Each step is independently committable and revertible. No node downtime is
required for steps 1–4 (watchdog and API changes take effect on next cycle).
## 7. Migration plan (Option A, if later authorized)
1. Human: purchase Cloudflare Zero Trust plan + 5 dedicated egress IPs.
2. Create the Zero Trust organization; create 5 egress policies, one per
node identity, each pinned to its dedicated egress IP.
3. On bl: enroll each node's Warp client into the org (replaces wgcf
consumer registration); per-node enrollment certs stored like current
confs (root-owned, 0600).
4. Rolling cutover, one node at a time: swap the node's WireGuard config to
the Zero Trust endpoint, run `netvm-node-up.sh <node>`, verify distinct
egress IP (see §8), then proceed.
5. Update NODES.md egress_ip column per node; commit.
## 8. Verification
Per-node, after any change:
```bash
for n in muse pip 646 opm def; do
ip=$(sudo ip netns exec warp-$n curl -s -m 10 https://api.ipify.org)
echo "$n: $ip"
done | tee /tmp/egress-check-$(date +%F).log
```
Acceptance for Option C:
- `HEALTH_FAIL_REASON="warp egress down"` appears in watchdog logs when a
tunnel is (test-)partitioned, and the node recovers without manual
intervention.
- Forcing a partition on ONE node (e.g. `wg set <WG> peer <PK>
remove` in a test netns) does not disturb the other nodes' automation.
- Fleet polling loops are measurably staggered in logs.
Acceptance for Option A:
- The 5 egress IPs are pairwise distinct AND stable across 24h
(hourly probe, logged).
- Rate-limiting one node's IP (if testable) leaves the others unaffected.
## 9. Rollback
- **Option C rollback:** revert the watchdog/API commits; the change is
purely additive to health checking — old behavior restored on next
watchdog cycle. No credential or network state to unwind.
- **Identity/config rollback:** `/root/netvm-backup-<date>/` holds the
pre-change `/etc/netvm/*.conf` copies; restore and re-run
`netvm-node-up.sh <node>` per node.
- **Option A rollback:** swap each node's WireGuard config back to the
wgcf consumer conf (kept in backup), re-run `netvm-node-up.sh`.
## 10. Open items (human-gated)
1. Option A requires a paid Cloudflare Zero Trust plan + dedicated egress
IP add-on — purchase decision belongs to the user.
2. If Cloudflare hard-blocks 104.28.195.181 (not just throttles), Option C
detects it but cannot route around it — that event should trigger the
Option A conversation.
3. `def` node is unassigned; include it in the rollout for uniformity, or
explicitly exclude it and note why.
+60
View File
@@ -0,0 +1,60 @@
# dm.py Cross-Operator DMs — Operator Runbook
## The `--to` Flag
Send messages between operators (opm, 646, pip, muse) using the recipient's
browser session. The sender is attributed via `[from <agent>]` prefix.
### Syntax
```
dm.py send --agent <sender> --to <recipient> --target <chat> "message"
```
- `--agent`: Sender identity (for `[from X]` attribution)
- `--to`: Recipient operator (whose browser/chat to post to)
- `--target`: `main` or a side chat name (must exist)
- Without `--to`: sends to own chat (legacy behavior)
### Examples
**opm → 646 (main chat):**
```
dm.py send --agent opm --to 646 --target main "hello from opm"
# Posts to 646's main as: [from opm] [uuid] hello from opm
```
**646 → opm (main chat):**
```
dm.py send --agent 646 --to opm --target main "hello from 646"
# Posts to opm's main as: [from 646] [uuid] hello from 646
```
**opm → 646 (side chat, must exist):**
```
dm.py send --agent opm --to 646 --target dm-opm "side thread msg"
# Posts to 646's 'dm-opm' side chat
```
### Reading
```
dm.py read --agent 646 --target main --n 5
# Reads 646's main chat (use recipient's --agent)
```
### Verified Working (2026-10-03)
- opm → 646 via main: ✓ [aea2809b]
- 646 → opm via main: ✓ [2d4fdb1d]
- Attribution `[from X]` appears correctly both ways
### Limitations
- **No `sidechat create`**: Target side chats must exist. Create via UI manually.
`muse-chat-api.py` has only `use`/`main`/`list` — no create API found.
- **Delivery NOT confirmed**: `SENT` means posted to DOM, not verified by recipient.
Confirm via independent read or recipient ack.
- **Backups**: `dm.py.bak` (pre---to), `dm.py.bak-20261003-toflag` (646's NameError fix)
### Design Notes
- DM via `dm.py` is the primary operator plane.
- Side chats preferred for direct calls (efficient, autonomous, threaded).
- Main chat for broadcasts or when side chat doesn't exist.
- Everything else (board, etc.) is fallback if DM goes down.