168 lines
10 KiB
Markdown
168 lines
10 KiB
Markdown
# NetVM
|
|
|
|
> [!IMPORTANT]
|
|
> **CRITICAL POLICY: MAIN CHAT PRESERVATION**
|
|
> Avoid using Main Chat whenever possible. When Main Chat gets bogged down with automated entries, scheduled job triggers, log dumps, or chatter, **work stops actually getting done**.
|
|
> All inputs, scheduled jobs, health checks, and inter-agent coordination must be relayed via designated **sidechats** / **side agents**, or provided via **file transfers**. See [CHAT_POLICY.md](file:///home/super/Projects/NetVM/CHAT_POLICY.md) for full specifications.
|
|
|
|
Fleet networking layer. Every node gets a stable network identity; every
|
|
byte of automation traffic is attributable, consistent, and boring — the
|
|
way good citizens look to the rest of the internet.
|
|
|
|
## Scope
|
|
|
|
1. **Per-node egress identity** — one Warp-backed egress per chrome-box /
|
|
node, in its own network namespace. Stable egress IP per node, explicit
|
|
mapping, isolated blast radius.
|
|
2. **Tailscale fabric** — node addressing, operator SSH, human reachability.
|
|
3. **Human-handoff links** — the reachable path a human uses to complete
|
|
CAPTCHA/2FA on a live browser session (original scope, kept).
|
|
|
|
NetVM owns networking. chrome-box owns browsers and consumes network from
|
|
NetVM. email-alert owns notifications. The Account(s) hub owns identity
|
|
records. Secrets and identities are the human's — agents manage structure
|
|
and lifecycle only, never credentials.
|
|
|
|
## Why per-node egress
|
|
|
|
- **Anti-fraud, not evasion.** Providers flag IP-hopping as suspicious. A
|
|
node that always egresses from the same place looks like what it is: a
|
|
legitimate machine. Standardizing on Warp is the good-citizen move.
|
|
- **Blast-radius isolation.** One IP reputation problem affects one node,
|
|
never the fleet. Nodes deliberately do not all share one egress.
|
|
- **DevOps standardization.** Every node is provisioned the same way:
|
|
netns + WireGuard-via-Warp + topology entry. No snowflakes.
|
|
|
|
Nodes may share an egress identity deliberately (documented in NODES.md);
|
|
sharing is a designed topology choice, never an accident.
|
|
|
|
## Architecture
|
|
|
|
chrome-box profile (client X)
|
|
-> launched inside netns warp-clientX (NetVM provides the launcher)
|
|
-> wg interface (Warp WireGuard params, human-generated config)
|
|
-> Cloudflare edge, stable colo (egress IP in NODES.md)
|
|
-> internet
|
|
|
|
operator / human
|
|
-> Tailscale tailnet (100.x) (management + Waypipe handoff)
|
|
-> node tail IP
|
|
|
|
Warp runs at the network layer, so the browser needs no proxy config —
|
|
the whole namespace egresses through Warp.
|
|
|
|
### Pattern decision
|
|
|
|
Two ways to get per-node Warp egress were considered:
|
|
|
|
- **warp-cli proxy mode** (mode proxy + SOCKS5 per node): simpler, but
|
|
warp-cli talks to a single system daemon (warp-svc) — running N daemons
|
|
on one host is unverified and fights the service model.
|
|
- **WireGuard in netns (chosen)**: Warp is WireGuard under the hood. One
|
|
wg interface per node namespace, config generated once by the human
|
|
(wgcf or equivalent), no daemon, no D-Bus, fully scriptable. One
|
|
interface per chrome-box, each independently up/down-able.
|
|
|
|
Upgrade path if pool IPs prove too fluid: Cloudflare Zero Trust dedicated
|
|
egress (true static IPs, paid).
|
|
|
|
## Mapping: 1:1 profile = node = warp identity = egress IP
|
|
|
|
Each chrome-box profile with auth gets its own dedicated Warp
|
|
credentials and its own stable egress IP. The profile name IS the node
|
|
name. One account always on one stable IP is the most human-like
|
|
pattern — it's IP hopping that trips provider alarms.
|
|
|
|
Interface names are hashed (`wb-<tag>`, `ve-<tag>`) because Linux
|
|
interface names max out at 15 chars; the netns keeps the full
|
|
`warp-<node>` name. See `bin/netvm-names.sh`.
|
|
|
|
## Provisioning a node
|
|
|
|
Identity creation is the human's job; lifecycle is scriptable:
|
|
|
|
1. Human: `chrome-box create <profile>` (browser profile).
|
|
2. Human: run `bin/netvm-new-identity.sh <profile>` ON the node — it
|
|
registers the Warp identity and installs /etc/netvm/<profile>.conf
|
|
(root-owned, 0600). This file is a credential — agents never create,
|
|
read, or copy it, and the script is excluded from the operator sudoers
|
|
allowlist.
|
|
3. Human (or operator over the tailnet): `bin/netvm-chrome.sh <profile> [url]`
|
|
— ensures the tunnel is up, enters netns `warp-<profile>`, bind-mounts
|
|
a working resolv.conf (the host's systemd-resolved stub is unreachable
|
|
in the netns), drops to the invoking user, launches the profile's
|
|
Chromium. Browser runs as the user, never as root.
|
|
|
|
Tear down: `bin/netvm-node-down.sh <node>` (sudo).
|
|
|
|
## Agent automation via CDP
|
|
|
|
Every profile launch exposes Chrome DevTools Protocol (deterministic port
|
|
per profile, see `bin/netvm-names.sh`). Chromium binds DevTools to loopback
|
|
only, so node-up runs a tiny TCP relay (`bin/netvm-cdp-relay.py`,
|
|
pidfile-supervised) from the veth IP to loopback.
|
|
|
|
- Local (on the node host): `curl http://<veth-ip>:<cdp-port>/json/list`
|
|
- Remote operator: `ssh -L <port>:<veth-ip>:<port> <user>@<tail-ip>` then
|
|
point Playwright (`connect_over_cdp`) or Puppeteer (`puppeteer.connect`)
|
|
at `http://127.0.0.1:<port>`.
|
|
- `bin/netvm-cdp.sh <profile>` prints the endpoint and the SSH command.
|
|
- `bin/netvm-chrome.sh --headless <profile>` runs without a display.
|
|
- CDP `Page.startScreencast` streams the live page — the human handoff path
|
|
(e.g. CAPTCHA): agent screencasts, human completes, agent resumes.
|
|
|
|
CDP is full browser control (cookies included). Exposure is host-local:
|
|
veth IPs aren't routable off the host and Warp forwards no inbound traffic.
|
|
|
|
## Files
|
|
|
|
- bin/netvm-verify.sh — prerequisite checks (ip, wg, netns, tailscale…).
|
|
- bin/netvm-node-up.sh <node> — bring up a node's egress.
|
|
- bin/netvm-node-down.sh <node> — tear a node's egress down.
|
|
- bin/netvm-topology.sh — print the live topology table.
|
|
- `bin/netvm-provision-edge.sh` — prepare an edge device (sudoers allowlist, deps, /etc/netvm); run on the node, once.
|
|
- `bin/netvm-fleet.sh` — operator fleet control over the tailnet (topology/up/down/ssh/exec/cdp per node).
|
|
- `bin/netvm-exec.sh <node> -- <cmd>` — run a command inside the node's netns as the invoking user (the agent-friendly primitive).
|
|
- `bin/netvm-enter.sh` — root worker behind netvm-exec/netvm-chrome (allowlisted; enters netns, fixes DNS, drops privs).
|
|
- `bin/netvm-chrome.sh <profile> [url]` — launch a chrome-box profile in its netns (CDP on by default; --headless for agents).
|
|
- `bin/netvm-cdp.sh <profile>` — print the CDP endpoint + SSH forward.
|
|
- `bin/netvm-cdp-relay.py` — veth-IP→loopback TCP relay for CDP (pidfile-supervised).
|
|
- `bin/netvm-names.sh` — shared naming: netns, hashed iface tags, veth subnet, CDP port.
|
|
- NODES.md — the network registry: profile/node -> netns -> Warp identity -> veth IP -> CDP port -> egress IP.
|
|
- ACCOUNTS.md — the secret-free login registry: login -> profile/node -> purpose -> auth state (no credentials, ever).
|
|
- `bin/netvm-accounts.sh` — operator view: registry joined with live node state.
|
|
- `bin/meta-ac-snapshot.py` — Meta Accounts Center change detector: CDP snapshot
|
|
(redirect chain + DOM markers) diffed against `snapshots/meta-ac/baseline.json`;
|
|
outcomes PASS/CHANGED/FAIL, `--promote` after human review.
|
|
- docs/META-ACCOUNTS-API.md — separate API for accountscenter.meta.com
|
|
(linkage/security surface; credential-isolated from phone-OTP).
|
|
- docs/PHONE-OTP.md — phone-number OTP login flow for muse.ai (proven on bl).
|
|
- `bin/accounts-health.py` — per-account CDP session probe (runs inside the netns).
|
|
- `bin/accounts-health.sh` — aggregates account vitality from ACCOUNTS.md,
|
|
signs + POSTs to the board health ingest (systemd timer, every 15 min).
|
|
- `bin/muse -a <account> [args]` — interactive terminal entrypoint for muse-cli; enforces account selection, auto-refreshes CDP cookies, and provides interactive email/OTP prompt fallback.
|
|
- `bin/muse-cli-node <node> [args]` — runs muse-cli inside node's netns with dedicated Cloudflare WARP egress & auto-refreshing cookies.
|
|
- `bin/refresh-node-cookies.py <node>` — extracts fresh cookies from running Chromium CDP in netns into `~/.config/muse-cli/<node>/cookies.txt`.
|
|
- `bin/muse_hybrid.py` — programmatic hybrid bridge combining fast gateway calls with CDP fallbacks.
|
|
- `bin/muse-tmux.py` — shared tmux socket manager (`/tmp/tmux-muse.sock`) for agent background execution, pipe-pane logging, and 2h session pruning.
|
|
- `bin/agent_md.py` — CLI & library for auditing, reading, writing, and synchronizing agent `.md` drive files (`SOUL.md`, `PROACTIVE_PREFERENCES.md`, `HEARTBEAT.md`, etc.) across containers via Hatch WebSocket RPC.
|
|
- `bin/agent-drive-watchdog.py` — background drive watchdog and auto-healing daemon (every 10m via `agent-drive-watchdog.timer`).
|
|
- `bin/swarm_worker/` & `bin/swarm-worker-supervise.sh` — supervised autonomous swarm worker daemon executing queued tasks in a hard sandbox.
|
|
- `bin/fleet-alert-relay.sh` — idempotent alert relay with 3-gate deduplication (watermark + 10m TTL hash + receipt verification) posting critical conditions to `#lobby`.
|
|
- `shared/operators/` — canonical operator drive markdown templates ensuring agents maintain autonomous loops, active supervision, and self-healing reflexes.
|
|
- docs/OPERATOR-DRIVE-RUNBOOK.md — operator runbook for auditing and modifying agent `.md` files via Hatch WebSocket RPC and SSH reverse tunnels.
|
|
- docs/HYBRID-GATEWAY-ADAPTATION.md — architectural guide on the muse-cli fast gateway adaptation and per-node egress isolation.
|
|
- docs/AGENT-TOOLING.md — guide to agent delegation, shared tmux background tooling, Work Orders (`[WO:...]`), and prompt envelope execution.
|
|
|
|
|
|
## Verification checklist
|
|
|
|
- [x] netvm-verify.sh passes on the target host. (tp, 2026-10-03)
|
|
- [x] Per-node wg interface comes up inside its netns; egress IP differs
|
|
per node (or matches the designed sharing in NODES.md). (tp: warp-tp, egress 104.28.203.246, 2026-10-03)
|
|
- [ ] Egress IP per node stable over 7 days (same colo pool).
|
|
- [ ] Chromium launched in the netns reaches the internet via the node's
|
|
egress (ip netns exec warp-<node> curl ifconfig.me).
|
|
- [ ] Waypipe handoff to a node over the tailnet shows a live browser.
|
|
- [ ] One node's egress killed -> other nodes unaffected (blast radius).
|