239 lines
11 KiB
Markdown
239 lines
11 KiB
Markdown
|
|
# WARP Egress Fix — Distinct Egress IPs per NetVM Node
|
|||
|
|
|
|||
|
|
Date: 2026-10-04
|
|||
|
|
Status: DESIGN (not yet implemented)
|
|||
|
|
Author: operator-main
|
|||
|
|
|
|||
|
|
## 1. Goal
|
|||
|
|
|
|||
|
|
Give each NetVM node (`muse`, `pip`, `646`, `opm`, `def`) its own distinct
|
|||
|
|
public egress IP, so that a Cloudflare rate-limit / block / partition on one
|
|||
|
|
node's egress IP does not take down the whole fleet at once.
|
|||
|
|
|
|||
|
|
Background: the 2026-10-04 partitioning investigation found all nodes egress
|
|||
|
|
from **one shared IP (104.28.195.181)**. Cloudflare sees the fleet as a single
|
|||
|
|
client; a single throttle/block correlates into a fleet-wide outage — which
|
|||
|
|
matches the observed "all 4 browsers flapping" pattern better than four
|
|||
|
|
independent failures.
|
|||
|
|
|
|||
|
|
## 2. Empirical finding (corrects the premise)
|
|||
|
|
|
|||
|
|
**Each node ALREADY has a distinct Warp identity** — and still shares one
|
|||
|
|
egress IP.
|
|||
|
|
|
|||
|
|
Verified live 2026-10-04 ~21:40 UTC on bl:
|
|||
|
|
|
|||
|
|
| node | identity (sha256 of PrivateKey, truncated) | egress IP now |
|
|||
|
|
|------|--------------------------------------------|---------------|
|
|||
|
|
| muse | 0fce50a5aff0 | 104.28.195.181 |
|
|||
|
|
| pip | e9fbb39cc4c5 | 104.28.195.181 |
|
|||
|
|
| 646 | e764d9d80fa8 | 104.28.195.181 |
|
|||
|
|
| opm | 771725bf8a5b | 104.28.195.181 |
|
|||
|
|
| def | fa868411c422 | 104.28.195.181 |
|
|||
|
|
|
|||
|
|
Identities live at `/etc/netvm/<node>.conf` (root-owned, 0600), one per node,
|
|||
|
|
generated via `netvm-new-identity.sh`. Egress IPs probed per-node with
|
|||
|
|
`ip netns exec warp-<node> curl -s https://api.ipify.org`.
|
|||
|
|
|
|||
|
|
**Why distinct identities share one egress IP:** consumer Cloudflare Warp
|
|||
|
|
(wgcf-registered) exits through Cloudflare's anycast edge. The egress IP is
|
|||
|
|
determined by the serving PoP/edge, not by the client identity — it is drawn
|
|||
|
|
from a pool shared across many Warp users. Registering more identities does
|
|||
|
|
not change the egress IP.
|
|||
|
|
|
|||
|
|
**Consequence:** "generate 4 distinct Warp identities" is already done and is
|
|||
|
|
NOT the fix. The fix must change *how traffic exits*, not *who the client is*.
|
|||
|
|
|
|||
|
|
## 3. Options
|
|||
|
|
|
|||
|
|
### Option A — Cloudflare Zero Trust dedicated egress IPs (paid, native)
|
|||
|
|
|
|||
|
|
Cloudflare Zero Trust (WARP for Teams) offers **dedicated egress IPs** as a
|
|||
|
|
paid add-on. With one dedicated egress IP per node and Gateway egress
|
|||
|
|
policies matched on device identity, each node gets a stable, distinct,
|
|||
|
|
organization-owned egress IP.
|
|||
|
|
|
|||
|
|
- Pros: stays on Cloudflare/WireGuard; `netvm-node-up.sh` flow mostly
|
|||
|
|
unchanged (different endpoint + enrollment cert); stable IPs you control;
|
|||
|
|
per-node egress policies give true isolation.
|
|||
|
|
- Cons: paid plan + per-IP add-on (human decision + payment); replaces
|
|||
|
|
wgcf consumer identities with Zero Trust enrollment — the largest config
|
|||
|
|
change of the four options; enrollment certs become new credentials to
|
|||
|
|
guard (same transient-handling rules as Warp keys).
|
|||
|
|
- Human-gated: yes (purchase + org setup).
|
|||
|
|
|
|||
|
|
### Option B — Per-node commercial VPN providers (distinct IPs, non-Warp)
|
|||
|
|
|
|||
|
|
Give each node a WireGuard config from a *different* commercial VPN provider
|
|||
|
|
(e.g. Mullvad, ProtonVPN, IVPN), each with its own exit IP.
|
|||
|
|
|
|||
|
|
- Pros: genuinely distinct egress IPs, guaranteed by construction; no
|
|||
|
|
Cloudflare correlation at all.
|
|||
|
|
- Cons: one account + credentials per provider (5 accounts); non-uniform
|
|||
|
|
configs; recurring cost per provider; abandons the Warp uniformity the
|
|||
|
|
fleet was built on; each provider's ToS / automation posture must be
|
|||
|
|
vetted.
|
|||
|
|
- Human-gated: yes (accounts + payment).
|
|||
|
|
|
|||
|
|
### Option C — Accept shared egress; de-correlate and detect (no new infra)
|
|||
|
|
|
|||
|
|
Keep consumer Warp. Fix the *failure mode* instead of the IP:
|
|||
|
|
|
|||
|
|
1. **Warp egress health stage in `chromebox-watchdog.sh`** (5th stage after
|
|||
|
|
the existing 4): check `wg show <WG> latest-handshakes` age AND probe
|
|||
|
|
egress from inside the netns
|
|||
|
|
(`ip netns exec warp-<node> curl -m5 https://www.cloudflare.com/cdn-cgi/trace`).
|
|||
|
|
On stale handshake or failed probe set
|
|||
|
|
`HEALTH_FAIL_REASON="warp egress down"` and take the same relaunch path
|
|||
|
|
as a dead CDP. This converts the invisible partition (CDP green, browser
|
|||
|
|
offline) into a visible, recoverable event.
|
|||
|
|
2. **React-readiness "match" assertion** in `muse-chat-api.py` before
|
|||
|
|
send/verify: assert title contains `Chat —` and the message-list
|
|||
|
|
container (`role="log"`) is present and non-skeleton. Fail closed
|
|||
|
|
otherwise. (A partitioned browser otherwise passes naive readiness.)
|
|||
|
|
3. **Per-node rate limiting + jitter** on muse.ai-bound traffic so the four
|
|||
|
|
nodes don't look like one coordinated client. `bin/rate_limiter.py`
|
|||
|
|
already exists — ensure token buckets are per-node, not shared, and add
|
|||
|
|
jitter to the 5-minute polling loops.
|
|||
|
|
4. **Stagger the fleet polling loops** (heartbeat, main-loop digests,
|
|||
|
|
health checks) so four nodes never fire in lockstep.
|
|||
|
|
|
|||
|
|
- Pros: no cost, no new credentials, no account changes; directly fixes the
|
|||
|
|
observed failure (invisible partition) and the most likely trigger
|
|||
|
|
(correlated rate-limiting); fully reversible.
|
|||
|
|
- Cons: egress IP stays shared — a hard Cloudflare block of 104.28.195.181
|
|||
|
|
still hits all nodes (but now it is *detected*, *logged distinctly*, and
|
|||
|
|
*recovered from* instead of flapping silently).
|
|||
|
|
- Human-gated: no.
|
|||
|
|
|
|||
|
|
### Option D — Hybrid: per-node egress proxies for sensitive paths only
|
|||
|
|
|
|||
|
|
Keep Warp for general traffic; route muse.ai-bound traffic per node through
|
|||
|
|
distinct SOCKS5/HTTP proxies with distinct IPs (browser `--proxy-server`
|
|||
|
|
per profile).
|
|||
|
|
|
|||
|
|
- Pros: distinct IPs where it matters (muse.ai automation); rest of stack
|
|||
|
|
untouched.
|
|||
|
|
- Cons: proxy accounts + credentials per node; added latency on the hot
|
|||
|
|
path; proxy reliability becomes a new failure domain; Chromium proxy
|
|||
|
|
config per profile adds bookkeeping.
|
|||
|
|
- Human-gated: partially (proxy accounts).
|
|||
|
|
|
|||
|
|
## 4. Recommendation
|
|||
|
|
|
|||
|
|
**Implement Option C now.** It is the only option that fixes the failure we
|
|||
|
|
actually observed (partitions invisible to the watchdog) with zero new cost
|
|||
|
|
or credentials, and it is fully reversible.
|
|||
|
|
|
|||
|
|
**Keep Option A as the paid path** if truly distinct, stable egress IPs are
|
|||
|
|
later required (e.g. Cloudflare starts hard-blocking the shared egress IP,
|
|||
|
|
or per-node IP reputation becomes a product need). Option A is the
|
|||
|
|
architecturally clean answer; it just needs a human purchase decision.
|
|||
|
|
|
|||
|
|
**Do not pursue Option B** unless the fleet leaves Cloudflare for other
|
|||
|
|
reasons — the ops overhead (5 VPN accounts) outweighs the benefit while
|
|||
|
|
Option C mitigates the correlation.
|
|||
|
|
|
|||
|
|
## 5. Identity rotation procedure (for reference)
|
|||
|
|
|
|||
|
|
Not needed for this fix (identities are already distinct), but recorded
|
|||
|
|
because the task asked how identities are generated. Per AGENTS.md, the user
|
|||
|
|
has explicitly authorized operators to run `netvm-new-identity.sh`
|
|||
|
|
("Warp credentials are NOT top secret").
|
|||
|
|
|
|||
|
|
One node at a time, rolling:
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
# on bl, as super (sudo -n via allowlist for netvm-node-up.sh)
|
|||
|
|
NODE=muse # repeat per node
|
|||
|
|
sudo cp /etc/netvm/${NODE}.conf /root/netvm-backup-$(date +%F)/${NODE}.conf
|
|||
|
|
sudo rm /etc/netvm/${NODE}.conf
|
|||
|
|
sudo /home/super/Projects/NetVM/bin/netvm-new-identity.sh ${NODE}
|
|||
|
|
sudo /home/super/Projects/NetVM/bin/netvm-node-up.sh ${NODE}
|
|||
|
|
ip netns exec warp-${NODE} curl -s -m 10 https://api.ipify.org; echo " <- ${NODE}"
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
`netvm-node-up.sh` is idempotent: it rebuilds the netns, veth pair,
|
|||
|
|
WireGuard interface, and CDP relay from the conf. Wait for the new tunnel's
|
|||
|
|
handshake (`wg show wb-<tag> latest-handshakes`) before moving to the next
|
|||
|
|
node. Never rotate two nodes simultaneously — keep 4/5 of the fleet up.
|
|||
|
|
|
|||
|
|
## 6. Migration plan (Option C)
|
|||
|
|
|
|||
|
|
Order matters: instrument first, then de-correlate.
|
|||
|
|
|
|||
|
|
1. **Watchdog 5th stage** — edit `bin/chromebox-watchdog.sh` `healthy()`:
|
|||
|
|
add Warp handshake-age check + in-netns egress probe; set
|
|||
|
|
`HEALTH_FAIL_REASON="warp egress down"`. Deploy to one node (opm) first,
|
|||
|
|
watch one full 5-minute cycle, then roll to all nodes. Commit.
|
|||
|
|
2. **React-readiness assertion** — add the match check in `muse-chat-api.py`
|
|||
|
|
before send/verify paths. Test against a known-good and a known-bad
|
|||
|
|
(partitioned or landing-page) browser. Commit.
|
|||
|
|
3. **Per-node rate limiter audit** — confirm `bin/rate_limiter.py` buckets
|
|||
|
|
are keyed per node; add jitter to fleet polling loops (heartbeat timer,
|
|||
|
|
self-main-loop.timer, health checks) so nodes don't fire in lockstep.
|
|||
|
|
Commit.
|
|||
|
|
4. **Stagger verification** — after deploy, confirm via logs that the four
|
|||
|
|
nodes' 5-minute loops are spread across the minute, not clustered.
|
|||
|
|
|
|||
|
|
Each step is independently committable and revertible. No node downtime is
|
|||
|
|
required for steps 1–4 (watchdog and API changes take effect on next cycle).
|
|||
|
|
|
|||
|
|
## 7. Migration plan (Option A, if later authorized)
|
|||
|
|
|
|||
|
|
1. Human: purchase Cloudflare Zero Trust plan + 5 dedicated egress IPs.
|
|||
|
|
2. Create the Zero Trust organization; create 5 egress policies, one per
|
|||
|
|
node identity, each pinned to its dedicated egress IP.
|
|||
|
|
3. On bl: enroll each node's Warp client into the org (replaces wgcf
|
|||
|
|
consumer registration); per-node enrollment certs stored like current
|
|||
|
|
confs (root-owned, 0600).
|
|||
|
|
4. Rolling cutover, one node at a time: swap the node's WireGuard config to
|
|||
|
|
the Zero Trust endpoint, run `netvm-node-up.sh <node>`, verify distinct
|
|||
|
|
egress IP (see §8), then proceed.
|
|||
|
|
5. Update NODES.md egress_ip column per node; commit.
|
|||
|
|
|
|||
|
|
## 8. Verification
|
|||
|
|
|
|||
|
|
Per-node, after any change:
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
for n in muse pip 646 opm def; do
|
|||
|
|
ip=$(sudo ip netns exec warp-$n curl -s -m 10 https://api.ipify.org)
|
|||
|
|
echo "$n: $ip"
|
|||
|
|
done | tee /tmp/egress-check-$(date +%F).log
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Acceptance for Option C:
|
|||
|
|
- `HEALTH_FAIL_REASON="warp egress down"` appears in watchdog logs when a
|
|||
|
|
tunnel is (test-)partitioned, and the node recovers without manual
|
|||
|
|
intervention.
|
|||
|
|
- Forcing a partition on ONE node (e.g. `wg set <WG> peer <PK>
|
|||
|
|
remove` in a test netns) does not disturb the other nodes' automation.
|
|||
|
|
- Fleet polling loops are measurably staggered in logs.
|
|||
|
|
|
|||
|
|
Acceptance for Option A:
|
|||
|
|
- The 5 egress IPs are pairwise distinct AND stable across 24h
|
|||
|
|
(hourly probe, logged).
|
|||
|
|
- Rate-limiting one node's IP (if testable) leaves the others unaffected.
|
|||
|
|
|
|||
|
|
## 9. Rollback
|
|||
|
|
|
|||
|
|
- **Option C rollback:** revert the watchdog/API commits; the change is
|
|||
|
|
purely additive to health checking — old behavior restored on next
|
|||
|
|
watchdog cycle. No credential or network state to unwind.
|
|||
|
|
- **Identity/config rollback:** `/root/netvm-backup-<date>/` holds the
|
|||
|
|
pre-change `/etc/netvm/*.conf` copies; restore and re-run
|
|||
|
|
`netvm-node-up.sh <node>` per node.
|
|||
|
|
- **Option A rollback:** swap each node's WireGuard config back to the
|
|||
|
|
wgcf consumer conf (kept in backup), re-run `netvm-node-up.sh`.
|
|||
|
|
|
|||
|
|
## 10. Open items (human-gated)
|
|||
|
|
|
|||
|
|
1. Option A requires a paid Cloudflare Zero Trust plan + dedicated egress
|
|||
|
|
IP add-on — purchase decision belongs to the user.
|
|||
|
|
2. If Cloudflare hard-blocks 104.28.195.181 (not just throttles), Option C
|
|||
|
|
detects it but cannot route around it — that event should trigger the
|
|||
|
|
Option A conversation.
|
|||
|
|
3. `def` node is unassigned; include it in the rollout for uniformity, or
|
|||
|
|
explicitly exclude it and note why.
|