# WARP Egress Fix — Distinct Egress IPs per NetVM Node Date: 2026-10-04 Status: DESIGN (not yet implemented) Author: operator-main ## 1. Goal Give each NetVM node (`muse`, `pip`, `646`, `opm`, `def`) its own distinct public egress IP, so that a Cloudflare rate-limit / block / partition on one node's egress IP does not take down the whole fleet at once. Background: the 2026-10-04 partitioning investigation found all nodes egress from **one shared IP (104.28.195.181)**. Cloudflare sees the fleet as a single client; a single throttle/block correlates into a fleet-wide outage — which matches the observed "all 4 browsers flapping" pattern better than four independent failures. ## 2. Empirical finding (corrects the premise) **Each node ALREADY has a distinct Warp identity** — and still shares one egress IP. Verified live 2026-10-04 ~21:40 UTC on bl: | node | identity (sha256 of PrivateKey, truncated) | egress IP now | |------|--------------------------------------------|---------------| | muse | 0fce50a5aff0 | 104.28.195.181 | | pip | e9fbb39cc4c5 | 104.28.195.181 | | 646 | e764d9d80fa8 | 104.28.195.181 | | opm | 771725bf8a5b | 104.28.195.181 | | def | fa868411c422 | 104.28.195.181 | Identities live at `/etc/netvm/.conf` (root-owned, 0600), one per node, generated via `netvm-new-identity.sh`. Egress IPs probed per-node with `ip netns exec warp- curl -s https://api.ipify.org`. **Why distinct identities share one egress IP:** consumer Cloudflare Warp (wgcf-registered) exits through Cloudflare's anycast edge. The egress IP is determined by the serving PoP/edge, not by the client identity — it is drawn from a pool shared across many Warp users. Registering more identities does not change the egress IP. **Consequence:** "generate 4 distinct Warp identities" is already done and is NOT the fix. The fix must change *how traffic exits*, not *who the client is*. ## 3. Options ### Option A — Cloudflare Zero Trust dedicated egress IPs (paid, native) Cloudflare Zero Trust (WARP for Teams) offers **dedicated egress IPs** as a paid add-on. With one dedicated egress IP per node and Gateway egress policies matched on device identity, each node gets a stable, distinct, organization-owned egress IP. - Pros: stays on Cloudflare/WireGuard; `netvm-node-up.sh` flow mostly unchanged (different endpoint + enrollment cert); stable IPs you control; per-node egress policies give true isolation. - Cons: paid plan + per-IP add-on (human decision + payment); replaces wgcf consumer identities with Zero Trust enrollment — the largest config change of the four options; enrollment certs become new credentials to guard (same transient-handling rules as Warp keys). - Human-gated: yes (purchase + org setup). ### Option B — Per-node commercial VPN providers (distinct IPs, non-Warp) Give each node a WireGuard config from a *different* commercial VPN provider (e.g. Mullvad, ProtonVPN, IVPN), each with its own exit IP. - Pros: genuinely distinct egress IPs, guaranteed by construction; no Cloudflare correlation at all. - Cons: one account + credentials per provider (5 accounts); non-uniform configs; recurring cost per provider; abandons the Warp uniformity the fleet was built on; each provider's ToS / automation posture must be vetted. - Human-gated: yes (accounts + payment). ### Option C — Accept shared egress; de-correlate and detect (no new infra) Keep consumer Warp. Fix the *failure mode* instead of the IP: 1. **Warp egress health stage in `chromebox-watchdog.sh`** (5th stage after the existing 4): check `wg show latest-handshakes` age AND probe egress from inside the netns (`ip netns exec warp- curl -m5 https://www.cloudflare.com/cdn-cgi/trace`). On stale handshake or failed probe set `HEALTH_FAIL_REASON="warp egress down"` and take the same relaunch path as a dead CDP. This converts the invisible partition (CDP green, browser offline) into a visible, recoverable event. 2. **React-readiness "match" assertion** in `muse-chat-api.py` before send/verify: assert title contains `Chat —` and the message-list container (`role="log"`) is present and non-skeleton. Fail closed otherwise. (A partitioned browser otherwise passes naive readiness.) 3. **Per-node rate limiting + jitter** on muse.ai-bound traffic so the four nodes don't look like one coordinated client. `bin/rate_limiter.py` already exists — ensure token buckets are per-node, not shared, and add jitter to the 5-minute polling loops. 4. **Stagger the fleet polling loops** (heartbeat, main-loop digests, health checks) so four nodes never fire in lockstep. - Pros: no cost, no new credentials, no account changes; directly fixes the observed failure (invisible partition) and the most likely trigger (correlated rate-limiting); fully reversible. - Cons: egress IP stays shared — a hard Cloudflare block of 104.28.195.181 still hits all nodes (but now it is *detected*, *logged distinctly*, and *recovered from* instead of flapping silently). - Human-gated: no. ### Option D — Hybrid: per-node egress proxies for sensitive paths only Keep Warp for general traffic; route muse.ai-bound traffic per node through distinct SOCKS5/HTTP proxies with distinct IPs (browser `--proxy-server` per profile). - Pros: distinct IPs where it matters (muse.ai automation); rest of stack untouched. - Cons: proxy accounts + credentials per node; added latency on the hot path; proxy reliability becomes a new failure domain; Chromium proxy config per profile adds bookkeeping. - Human-gated: partially (proxy accounts). ## 4. Recommendation **Implement Option C now.** It is the only option that fixes the failure we actually observed (partitions invisible to the watchdog) with zero new cost or credentials, and it is fully reversible. **Keep Option A as the paid path** if truly distinct, stable egress IPs are later required (e.g. Cloudflare starts hard-blocking the shared egress IP, or per-node IP reputation becomes a product need). Option A is the architecturally clean answer; it just needs a human purchase decision. **Do not pursue Option B** unless the fleet leaves Cloudflare for other reasons — the ops overhead (5 VPN accounts) outweighs the benefit while Option C mitigates the correlation. ## 5. Identity rotation procedure (for reference) Not needed for this fix (identities are already distinct), but recorded because the task asked how identities are generated. Per AGENTS.md, the user has explicitly authorized operators to run `netvm-new-identity.sh` ("Warp credentials are NOT top secret"). One node at a time, rolling: ```bash # on bl, as super (sudo -n via allowlist for netvm-node-up.sh) NODE=muse # repeat per node sudo cp /etc/netvm/${NODE}.conf /root/netvm-backup-$(date +%F)/${NODE}.conf sudo rm /etc/netvm/${NODE}.conf sudo /home/super/Projects/NetVM/bin/netvm-new-identity.sh ${NODE} sudo /home/super/Projects/NetVM/bin/netvm-node-up.sh ${NODE} ip netns exec warp-${NODE} curl -s -m 10 https://api.ipify.org; echo " <- ${NODE}" ``` `netvm-node-up.sh` is idempotent: it rebuilds the netns, veth pair, WireGuard interface, and CDP relay from the conf. Wait for the new tunnel's handshake (`wg show wb- latest-handshakes`) before moving to the next node. Never rotate two nodes simultaneously — keep 4/5 of the fleet up. ## 6. Migration plan (Option C) Order matters: instrument first, then de-correlate. 1. **Watchdog 5th stage** — edit `bin/chromebox-watchdog.sh` `healthy()`: add Warp handshake-age check + in-netns egress probe; set `HEALTH_FAIL_REASON="warp egress down"`. Deploy to one node (opm) first, watch one full 5-minute cycle, then roll to all nodes. Commit. 2. **React-readiness assertion** — add the match check in `muse-chat-api.py` before send/verify paths. Test against a known-good and a known-bad (partitioned or landing-page) browser. Commit. 3. **Per-node rate limiter audit** — confirm `bin/rate_limiter.py` buckets are keyed per node; add jitter to fleet polling loops (heartbeat timer, self-main-loop.timer, health checks) so nodes don't fire in lockstep. Commit. 4. **Stagger verification** — after deploy, confirm via logs that the four nodes' 5-minute loops are spread across the minute, not clustered. Each step is independently committable and revertible. No node downtime is required for steps 1–4 (watchdog and API changes take effect on next cycle). ## 7. Migration plan (Option A, if later authorized) 1. Human: purchase Cloudflare Zero Trust plan + 5 dedicated egress IPs. 2. Create the Zero Trust organization; create 5 egress policies, one per node identity, each pinned to its dedicated egress IP. 3. On bl: enroll each node's Warp client into the org (replaces wgcf consumer registration); per-node enrollment certs stored like current confs (root-owned, 0600). 4. Rolling cutover, one node at a time: swap the node's WireGuard config to the Zero Trust endpoint, run `netvm-node-up.sh `, verify distinct egress IP (see §8), then proceed. 5. Update NODES.md egress_ip column per node; commit. ## 8. Verification Per-node, after any change: ```bash for n in muse pip 646 opm def; do ip=$(sudo ip netns exec warp-$n curl -s -m 10 https://api.ipify.org) echo "$n: $ip" done | tee /tmp/egress-check-$(date +%F).log ``` Acceptance for Option C: - `HEALTH_FAIL_REASON="warp egress down"` appears in watchdog logs when a tunnel is (test-)partitioned, and the node recovers without manual intervention. - Forcing a partition on ONE node (e.g. `wg set peer remove` in a test netns) does not disturb the other nodes' automation. - Fleet polling loops are measurably staggered in logs. Acceptance for Option A: - The 5 egress IPs are pairwise distinct AND stable across 24h (hourly probe, logged). - Rate-limiting one node's IP (if testable) leaves the others unaffected. ## 9. Rollback - **Option C rollback:** revert the watchdog/API commits; the change is purely additive to health checking — old behavior restored on next watchdog cycle. No credential or network state to unwind. - **Identity/config rollback:** `/root/netvm-backup-/` holds the pre-change `/etc/netvm/*.conf` copies; restore and re-run `netvm-node-up.sh ` per node. - **Option A rollback:** swap each node's WireGuard config back to the wgcf consumer conf (kept in backup), re-run `netvm-node-up.sh`. ## 10. Open items (human-gated) 1. Option A requires a paid Cloudflare Zero Trust plan + dedicated egress IP add-on — purchase decision belongs to the user. 2. If Cloudflare hard-blocks 104.28.195.181 (not just throttles), Option C detects it but cannot route around it — that event should trigger the Option A conversation. 3. `def` node is unassigned; include it in the rollout for uniformity, or explicitly exclude it and note why.