Files
box/docs/WARP-EGRESS-FIX.md
T

11 KiB
Raw Blame History

WARP Egress Fix — Distinct Egress IPs per NetVM Node

Box is the main surface. All operator work goes through Box (box.muse-dev.online). The web UI, box CLI, and agents share the same API endpoints. No UI-only powers.

Date: 2026-10-04 Status: DESIGN (not yet implemented) Author: operator-main

1. Goal

Give each NetVM node (muse, pip, 646, opm, def) its own distinct public egress IP, so that a Cloudflare rate-limit / block / partition on one node's egress IP does not take down the whole fleet at once.

Background: the 2026-10-04 partitioning investigation found all nodes egress from one shared IP (104.28.195.181). Cloudflare sees the fleet as a single client; a single throttle/block correlates into a fleet-wide outage — which matches the observed "all 4 browsers flapping" pattern better than four independent failures.

2. Empirical finding (corrects the premise)

Each node ALREADY has a distinct Warp identity — and still shares one egress IP.

Verified live 2026-10-04 ~21:40 UTC on bl:

node identity (sha256 of PrivateKey, truncated) egress IP now
muse 0fce50a5aff0 104.28.195.181
pip e9fbb39cc4c5 104.28.195.181
646 e764d9d80fa8 104.28.195.181
opm 771725bf8a5b 104.28.195.181
def fa868411c422 104.28.195.181

Identities live at /etc/netvm/<node>.conf (root-owned, 0600), one per node, generated via netvm-new-identity.sh. Egress IPs probed per-node with ip netns exec warp-<node> curl -s https://api.ipify.org.

Why distinct identities share one egress IP: consumer Cloudflare Warp (wgcf-registered) exits through Cloudflare's anycast edge. The egress IP is determined by the serving PoP/edge, not by the client identity — it is drawn from a pool shared across many Warp users. Registering more identities does not change the egress IP.

Consequence: "generate 4 distinct Warp identities" is already done and is NOT the fix. The fix must change how traffic exits, not who the client is.

3. Options

Option A — Cloudflare Zero Trust dedicated egress IPs (paid, native)

Cloudflare Zero Trust (WARP for Teams) offers dedicated egress IPs as a paid add-on. With one dedicated egress IP per node and Gateway egress policies matched on device identity, each node gets a stable, distinct, organization-owned egress IP.

  • Pros: stays on Cloudflare/WireGuard; netvm-node-up.sh flow mostly unchanged (different endpoint + enrollment cert); stable IPs you control; per-node egress policies give true isolation.
  • Cons: paid plan + per-IP add-on (human decision + payment); replaces wgcf consumer identities with Zero Trust enrollment — the largest config change of the four options; enrollment certs become new credentials to guard (same transient-handling rules as Warp keys).
  • Human-gated: yes (purchase + org setup).

Option B — Per-node commercial VPN providers (distinct IPs, non-Warp)

Give each node a WireGuard config from a different commercial VPN provider (e.g. Mullvad, ProtonVPN, IVPN), each with its own exit IP.

  • Pros: genuinely distinct egress IPs, guaranteed by construction; no Cloudflare correlation at all.
  • Cons: one account + credentials per provider (5 accounts); non-uniform configs; recurring cost per provider; abandons the Warp uniformity the fleet was built on; each provider's ToS / automation posture must be vetted.
  • Human-gated: yes (accounts + payment).

Option C — Accept shared egress; de-correlate and detect (no new infra)

Keep consumer Warp. Fix the failure mode instead of the IP:

  1. Warp egress health stage in chromebox-watchdog.sh (5th stage after the existing 4): check wg show <WG> latest-handshakes age AND probe egress from inside the netns (ip netns exec warp-<node> curl -m5 https://www.cloudflare.com/cdn-cgi/trace). On stale handshake or failed probe set HEALTH_FAIL_REASON="warp egress down" and take the same relaunch path as a dead CDP. This converts the invisible partition (CDP green, browser offline) into a visible, recoverable event.
  2. React-readiness "match" assertion in muse-chat-api.py before send/verify: assert title contains Chat — and the message-list container (role="log") is present and non-skeleton. Fail closed otherwise. (A partitioned browser otherwise passes naive readiness.)
  3. Per-node rate limiting + jitter on muse.ai-bound traffic so the four nodes don't look like one coordinated client. bin/rate_limiter.py already exists — ensure token buckets are per-node, not shared, and add jitter to the 5-minute polling loops.
  4. Stagger the fleet polling loops (heartbeat, main-loop digests, health checks) so four nodes never fire in lockstep.
  • Pros: no cost, no new credentials, no account changes; directly fixes the observed failure (invisible partition) and the most likely trigger (correlated rate-limiting); fully reversible.
  • Cons: egress IP stays shared — a hard Cloudflare block of 104.28.195.181 still hits all nodes (but now it is detected, logged distinctly, and recovered from instead of flapping silently).
  • Human-gated: no.

Option D — Hybrid: per-node egress proxies for sensitive paths only

Keep Warp for general traffic; route muse.ai-bound traffic per node through distinct SOCKS5/HTTP proxies with distinct IPs (browser --proxy-server per profile).

  • Pros: distinct IPs where it matters (muse.ai automation); rest of stack untouched.
  • Cons: proxy accounts + credentials per node; added latency on the hot path; proxy reliability becomes a new failure domain; Chromium proxy config per profile adds bookkeeping.
  • Human-gated: partially (proxy accounts).

4. Recommendation

Implement Option C now. It is the only option that fixes the failure we actually observed (partitions invisible to the watchdog) with zero new cost or credentials, and it is fully reversible.

Keep Option A as the paid path if truly distinct, stable egress IPs are later required (e.g. Cloudflare starts hard-blocking the shared egress IP, or per-node IP reputation becomes a product need). Option A is the architecturally clean answer; it just needs a human purchase decision.

Do not pursue Option B unless the fleet leaves Cloudflare for other reasons — the ops overhead (5 VPN accounts) outweighs the benefit while Option C mitigates the correlation.

5. Identity rotation procedure (for reference)

Not needed for this fix (identities are already distinct), but recorded because the task asked how identities are generated. Per AGENTS.md, the user has explicitly authorized operators to run netvm-new-identity.sh ("Warp credentials are NOT top secret").

One node at a time, rolling:

# on bl, as super (sudo -n via allowlist for netvm-node-up.sh)
NODE=muse  # repeat per node
sudo cp /etc/netvm/${NODE}.conf /root/netvm-backup-$(date +%F)/${NODE}.conf
sudo rm /etc/netvm/${NODE}.conf
sudo /home/super/Projects/NetVM/bin/netvm-new-identity.sh ${NODE}
sudo /home/super/Projects/NetVM/bin/netvm-node-up.sh ${NODE}
ip netns exec warp-${NODE} curl -s -m 10 https://api.ipify.org; echo " <- ${NODE}"

netvm-node-up.sh is idempotent: it rebuilds the netns, veth pair, WireGuard interface, and CDP relay from the conf. Wait for the new tunnel's handshake (wg show wb-<tag> latest-handshakes) before moving to the next node. Never rotate two nodes simultaneously — keep 4/5 of the fleet up.

6. Migration plan (Option C)

Order matters: instrument first, then de-correlate.

  1. Watchdog 5th stage — edit bin/chromebox-watchdog.sh healthy(): add Warp handshake-age check + in-netns egress probe; set HEALTH_FAIL_REASON="warp egress down". Deploy to one node (opm) first, watch one full 5-minute cycle, then roll to all nodes. Commit.
  2. React-readiness assertion — add the match check in muse-chat-api.py before send/verify paths. Test against a known-good and a known-bad (partitioned or landing-page) browser. Commit.
  3. Per-node rate limiter audit — confirm bin/rate_limiter.py buckets are keyed per node; add jitter to fleet polling loops (heartbeat timer, self-main-loop.timer, health checks) so nodes don't fire in lockstep. Commit.
  4. Stagger verification — after deploy, confirm via logs that the four nodes' 5-minute loops are spread across the minute, not clustered.

Each step is independently committable and revertible. No node downtime is required for steps 1–4 (watchdog and API changes take effect on next cycle).

7. Migration plan (Option A, if later authorized)

  1. Human: purchase Cloudflare Zero Trust plan + 5 dedicated egress IPs.
  2. Create the Zero Trust organization; create 5 egress policies, one per node identity, each pinned to its dedicated egress IP.
  3. On bl: enroll each node's Warp client into the org (replaces wgcf consumer registration); per-node enrollment certs stored like current confs (root-owned, 0600).
  4. Rolling cutover, one node at a time: swap the node's WireGuard config to the Zero Trust endpoint, run netvm-node-up.sh <node>, verify distinct egress IP (see §8), then proceed.
  5. Update NODES.md egress_ip column per node; commit.

8. Verification

Per-node, after any change:

for n in muse pip 646 opm def; do
  ip=$(sudo ip netns exec warp-$n curl -s -m 10 https://api.ipify.org)
  echo "$n: $ip"
done | tee /tmp/egress-check-$(date +%F).log

Acceptance for Option C:

  • HEALTH_FAIL_REASON="warp egress down" appears in watchdog logs when a tunnel is (test-)partitioned, and the node recovers without manual intervention.
  • Forcing a partition on ONE node (e.g. wg set <WG> peer <PK> remove in a test netns) does not disturb the other nodes' automation.
  • Fleet polling loops are measurably staggered in logs.

Acceptance for Option A:

  • The 5 egress IPs are pairwise distinct AND stable across 24h (hourly probe, logged).
  • Rate-limiting one node's IP (if testable) leaves the others unaffected.

9. Rollback

  • Option C rollback: revert the watchdog/API commits; the change is purely additive to health checking — old behavior restored on next watchdog cycle. No credential or network state to unwind.
  • Identity/config rollback: /root/netvm-backup-<date>/ holds the pre-change /etc/netvm/*.conf copies; restore and re-run netvm-node-up.sh <node> per node.
  • Option A rollback: swap each node's WireGuard config back to the wgcf consumer conf (kept in backup), re-run netvm-node-up.sh.

10. Open items (human-gated)

  1. Option A requires a paid Cloudflare Zero Trust plan + dedicated egress IP add-on — purchase decision belongs to the user.
  2. If Cloudflare hard-blocks 104.28.195.181 (not just throttles), Option C detects it but cannot route around it — that event should trigger the Option A conversation.
  3. def node is unassigned; include it in the rollout for uniformity, or explicitly exclude it and note why.