239 lines
11 KiB
Markdown
239 lines
11 KiB
Markdown
# WARP Egress Fix — Distinct Egress IPs per NetVM Node
|
||
|
||
Date: 2026-10-04
|
||
Status: DESIGN (not yet implemented)
|
||
Author: operator-main
|
||
|
||
## 1. Goal
|
||
|
||
Give each NetVM node (`muse`, `pip`, `646`, `opm`, `def`) its own distinct
|
||
public egress IP, so that a Cloudflare rate-limit / block / partition on one
|
||
node's egress IP does not take down the whole fleet at once.
|
||
|
||
Background: the 2026-10-04 partitioning investigation found all nodes egress
|
||
from **one shared IP (104.28.195.181)**. Cloudflare sees the fleet as a single
|
||
client; a single throttle/block correlates into a fleet-wide outage — which
|
||
matches the observed "all 4 browsers flapping" pattern better than four
|
||
independent failures.
|
||
|
||
## 2. Empirical finding (corrects the premise)
|
||
|
||
**Each node ALREADY has a distinct Warp identity** — and still shares one
|
||
egress IP.
|
||
|
||
Verified live 2026-10-04 ~21:40 UTC on bl:
|
||
|
||
| node | identity (sha256 of PrivateKey, truncated) | egress IP now |
|
||
|------|--------------------------------------------|---------------|
|
||
| muse | 0fce50a5aff0 | 104.28.195.181 |
|
||
| pip | e9fbb39cc4c5 | 104.28.195.181 |
|
||
| 646 | e764d9d80fa8 | 104.28.195.181 |
|
||
| opm | 771725bf8a5b | 104.28.195.181 |
|
||
| def | fa868411c422 | 104.28.195.181 |
|
||
|
||
Identities live at `/etc/netvm/<node>.conf` (root-owned, 0600), one per node,
|
||
generated via `netvm-new-identity.sh`. Egress IPs probed per-node with
|
||
`ip netns exec warp-<node> curl -s https://api.ipify.org`.
|
||
|
||
**Why distinct identities share one egress IP:** consumer Cloudflare Warp
|
||
(wgcf-registered) exits through Cloudflare's anycast edge. The egress IP is
|
||
determined by the serving PoP/edge, not by the client identity — it is drawn
|
||
from a pool shared across many Warp users. Registering more identities does
|
||
not change the egress IP.
|
||
|
||
**Consequence:** "generate 4 distinct Warp identities" is already done and is
|
||
NOT the fix. The fix must change *how traffic exits*, not *who the client is*.
|
||
|
||
## 3. Options
|
||
|
||
### Option A — Cloudflare Zero Trust dedicated egress IPs (paid, native)
|
||
|
||
Cloudflare Zero Trust (WARP for Teams) offers **dedicated egress IPs** as a
|
||
paid add-on. With one dedicated egress IP per node and Gateway egress
|
||
policies matched on device identity, each node gets a stable, distinct,
|
||
organization-owned egress IP.
|
||
|
||
- Pros: stays on Cloudflare/WireGuard; `netvm-node-up.sh` flow mostly
|
||
unchanged (different endpoint + enrollment cert); stable IPs you control;
|
||
per-node egress policies give true isolation.
|
||
- Cons: paid plan + per-IP add-on (human decision + payment); replaces
|
||
wgcf consumer identities with Zero Trust enrollment — the largest config
|
||
change of the four options; enrollment certs become new credentials to
|
||
guard (same transient-handling rules as Warp keys).
|
||
- Human-gated: yes (purchase + org setup).
|
||
|
||
### Option B — Per-node commercial VPN providers (distinct IPs, non-Warp)
|
||
|
||
Give each node a WireGuard config from a *different* commercial VPN provider
|
||
(e.g. Mullvad, ProtonVPN, IVPN), each with its own exit IP.
|
||
|
||
- Pros: genuinely distinct egress IPs, guaranteed by construction; no
|
||
Cloudflare correlation at all.
|
||
- Cons: one account + credentials per provider (5 accounts); non-uniform
|
||
configs; recurring cost per provider; abandons the Warp uniformity the
|
||
fleet was built on; each provider's ToS / automation posture must be
|
||
vetted.
|
||
- Human-gated: yes (accounts + payment).
|
||
|
||
### Option C — Accept shared egress; de-correlate and detect (no new infra)
|
||
|
||
Keep consumer Warp. Fix the *failure mode* instead of the IP:
|
||
|
||
1. **Warp egress health stage in `chromebox-watchdog.sh`** (5th stage after
|
||
the existing 4): check `wg show <WG> latest-handshakes` age AND probe
|
||
egress from inside the netns
|
||
(`ip netns exec warp-<node> curl -m5 https://www.cloudflare.com/cdn-cgi/trace`).
|
||
On stale handshake or failed probe set
|
||
`HEALTH_FAIL_REASON="warp egress down"` and take the same relaunch path
|
||
as a dead CDP. This converts the invisible partition (CDP green, browser
|
||
offline) into a visible, recoverable event.
|
||
2. **React-readiness "match" assertion** in `muse-chat-api.py` before
|
||
send/verify: assert title contains `Chat —` and the message-list
|
||
container (`role="log"`) is present and non-skeleton. Fail closed
|
||
otherwise. (A partitioned browser otherwise passes naive readiness.)
|
||
3. **Per-node rate limiting + jitter** on muse.ai-bound traffic so the four
|
||
nodes don't look like one coordinated client. `bin/rate_limiter.py`
|
||
already exists — ensure token buckets are per-node, not shared, and add
|
||
jitter to the 5-minute polling loops.
|
||
4. **Stagger the fleet polling loops** (heartbeat, main-loop digests,
|
||
health checks) so four nodes never fire in lockstep.
|
||
|
||
- Pros: no cost, no new credentials, no account changes; directly fixes the
|
||
observed failure (invisible partition) and the most likely trigger
|
||
(correlated rate-limiting); fully reversible.
|
||
- Cons: egress IP stays shared — a hard Cloudflare block of 104.28.195.181
|
||
still hits all nodes (but now it is *detected*, *logged distinctly*, and
|
||
*recovered from* instead of flapping silently).
|
||
- Human-gated: no.
|
||
|
||
### Option D — Hybrid: per-node egress proxies for sensitive paths only
|
||
|
||
Keep Warp for general traffic; route muse.ai-bound traffic per node through
|
||
distinct SOCKS5/HTTP proxies with distinct IPs (browser `--proxy-server`
|
||
per profile).
|
||
|
||
- Pros: distinct IPs where it matters (muse.ai automation); rest of stack
|
||
untouched.
|
||
- Cons: proxy accounts + credentials per node; added latency on the hot
|
||
path; proxy reliability becomes a new failure domain; Chromium proxy
|
||
config per profile adds bookkeeping.
|
||
- Human-gated: partially (proxy accounts).
|
||
|
||
## 4. Recommendation
|
||
|
||
**Implement Option C now.** It is the only option that fixes the failure we
|
||
actually observed (partitions invisible to the watchdog) with zero new cost
|
||
or credentials, and it is fully reversible.
|
||
|
||
**Keep Option A as the paid path** if truly distinct, stable egress IPs are
|
||
later required (e.g. Cloudflare starts hard-blocking the shared egress IP,
|
||
or per-node IP reputation becomes a product need). Option A is the
|
||
architecturally clean answer; it just needs a human purchase decision.
|
||
|
||
**Do not pursue Option B** unless the fleet leaves Cloudflare for other
|
||
reasons — the ops overhead (5 VPN accounts) outweighs the benefit while
|
||
Option C mitigates the correlation.
|
||
|
||
## 5. Identity rotation procedure (for reference)
|
||
|
||
Not needed for this fix (identities are already distinct), but recorded
|
||
because the task asked how identities are generated. Per AGENTS.md, the user
|
||
has explicitly authorized operators to run `netvm-new-identity.sh`
|
||
("Warp credentials are NOT top secret").
|
||
|
||
One node at a time, rolling:
|
||
|
||
```bash
|
||
# on bl, as super (sudo -n via allowlist for netvm-node-up.sh)
|
||
NODE=muse # repeat per node
|
||
sudo cp /etc/netvm/${NODE}.conf /root/netvm-backup-$(date +%F)/${NODE}.conf
|
||
sudo rm /etc/netvm/${NODE}.conf
|
||
sudo /home/super/Projects/NetVM/bin/netvm-new-identity.sh ${NODE}
|
||
sudo /home/super/Projects/NetVM/bin/netvm-node-up.sh ${NODE}
|
||
ip netns exec warp-${NODE} curl -s -m 10 https://api.ipify.org; echo " <- ${NODE}"
|
||
```
|
||
|
||
`netvm-node-up.sh` is idempotent: it rebuilds the netns, veth pair,
|
||
WireGuard interface, and CDP relay from the conf. Wait for the new tunnel's
|
||
handshake (`wg show wb-<tag> latest-handshakes`) before moving to the next
|
||
node. Never rotate two nodes simultaneously — keep 4/5 of the fleet up.
|
||
|
||
## 6. Migration plan (Option C)
|
||
|
||
Order matters: instrument first, then de-correlate.
|
||
|
||
1. **Watchdog 5th stage** — edit `bin/chromebox-watchdog.sh` `healthy()`:
|
||
add Warp handshake-age check + in-netns egress probe; set
|
||
`HEALTH_FAIL_REASON="warp egress down"`. Deploy to one node (opm) first,
|
||
watch one full 5-minute cycle, then roll to all nodes. Commit.
|
||
2. **React-readiness assertion** — add the match check in `muse-chat-api.py`
|
||
before send/verify paths. Test against a known-good and a known-bad
|
||
(partitioned or landing-page) browser. Commit.
|
||
3. **Per-node rate limiter audit** — confirm `bin/rate_limiter.py` buckets
|
||
are keyed per node; add jitter to fleet polling loops (heartbeat timer,
|
||
self-main-loop.timer, health checks) so nodes don't fire in lockstep.
|
||
Commit.
|
||
4. **Stagger verification** — after deploy, confirm via logs that the four
|
||
nodes' 5-minute loops are spread across the minute, not clustered.
|
||
|
||
Each step is independently committable and revertible. No node downtime is
|
||
required for steps 1–4 (watchdog and API changes take effect on next cycle).
|
||
|
||
## 7. Migration plan (Option A, if later authorized)
|
||
|
||
1. Human: purchase Cloudflare Zero Trust plan + 5 dedicated egress IPs.
|
||
2. Create the Zero Trust organization; create 5 egress policies, one per
|
||
node identity, each pinned to its dedicated egress IP.
|
||
3. On bl: enroll each node's Warp client into the org (replaces wgcf
|
||
consumer registration); per-node enrollment certs stored like current
|
||
confs (root-owned, 0600).
|
||
4. Rolling cutover, one node at a time: swap the node's WireGuard config to
|
||
the Zero Trust endpoint, run `netvm-node-up.sh <node>`, verify distinct
|
||
egress IP (see §8), then proceed.
|
||
5. Update NODES.md egress_ip column per node; commit.
|
||
|
||
## 8. Verification
|
||
|
||
Per-node, after any change:
|
||
|
||
```bash
|
||
for n in muse pip 646 opm def; do
|
||
ip=$(sudo ip netns exec warp-$n curl -s -m 10 https://api.ipify.org)
|
||
echo "$n: $ip"
|
||
done | tee /tmp/egress-check-$(date +%F).log
|
||
```
|
||
|
||
Acceptance for Option C:
|
||
- `HEALTH_FAIL_REASON="warp egress down"` appears in watchdog logs when a
|
||
tunnel is (test-)partitioned, and the node recovers without manual
|
||
intervention.
|
||
- Forcing a partition on ONE node (e.g. `wg set <WG> peer <PK>
|
||
remove` in a test netns) does not disturb the other nodes' automation.
|
||
- Fleet polling loops are measurably staggered in logs.
|
||
|
||
Acceptance for Option A:
|
||
- The 5 egress IPs are pairwise distinct AND stable across 24h
|
||
(hourly probe, logged).
|
||
- Rate-limiting one node's IP (if testable) leaves the others unaffected.
|
||
|
||
## 9. Rollback
|
||
|
||
- **Option C rollback:** revert the watchdog/API commits; the change is
|
||
purely additive to health checking — old behavior restored on next
|
||
watchdog cycle. No credential or network state to unwind.
|
||
- **Identity/config rollback:** `/root/netvm-backup-<date>/` holds the
|
||
pre-change `/etc/netvm/*.conf` copies; restore and re-run
|
||
`netvm-node-up.sh <node>` per node.
|
||
- **Option A rollback:** swap each node's WireGuard config back to the
|
||
wgcf consumer conf (kept in backup), re-run `netvm-node-up.sh`.
|
||
|
||
## 10. Open items (human-gated)
|
||
|
||
1. Option A requires a paid Cloudflare Zero Trust plan + dedicated egress
|
||
IP add-on — purchase decision belongs to the user.
|
||
2. If Cloudflare hard-blocks 104.28.195.181 (not just throttles), Option C
|
||
detects it but cannot route around it — that event should trigger the
|
||
Option A conversation.
|
||
3. `def` node is unassigned; include it in the rollout for uniformity, or
|
||
explicitly exclude it and note why.
|