2026-10-04 22:54:01 +00:00
# WARP Egress Fix — Distinct Egress IPs per NetVM Node
2026-10-05 15:58:37 +00:00
> **Box is the main surface.** All operator work goes through Box (box.muse-dev.online). The web UI, `box` CLI, and agents share the same API endpoints. No UI-only powers.
2026-10-04 22:54:01 +00:00
Date: 2026-10-04
Status: DESIGN (not yet implemented)
Author: operator-main
## 1. Goal
Give each NetVM node (`muse` , `pip` , `646` , `opm` , `def` ) its own distinct
public egress IP, so that a Cloudflare rate-limit / block / partition on one
node's egress IP does not take down the whole fleet at once.
Background: the 2026-10-04 partitioning investigation found all nodes egress
from **one shared IP (104.28.195.181) ** . Cloudflare sees the fleet as a single
client; a single throttle/block correlates into a fleet-wide outage — which
matches the observed "all 4 browsers flapping" pattern better than four
independent failures.
## 2. Empirical finding (corrects the premise)
**Each node ALREADY has a distinct Warp identity ** — and still shares one
egress IP.
Verified live 2026-10-04 ~21:40 UTC on bl:
| node | identity (sha256 of PrivateKey, truncated) | egress IP now |
|------|--------------------------------------------|---------------|
| muse | 0fce50a5aff0 | 104.28.195.181 |
| pip | e9fbb39cc4c5 | 104.28.195.181 |
| 646 | e764d9d80fa8 | 104.28.195.181 |
| opm | 771725bf8a5b | 104.28.195.181 |
| def | fa868411c422 | 104.28.195.181 |
Identities live at `/etc/netvm/<node>.conf` (root-owned, 0600), one per node,
generated via `netvm-new-identity.sh` . Egress IPs probed per-node with
`ip netns exec warp-<node> curl -s https://api.ipify.org` .
**Why distinct identities share one egress IP: ** consumer Cloudflare Warp
(wgcf-registered) exits through Cloudflare's anycast edge. The egress IP is
determined by the serving PoP/edge, not by the client identity — it is drawn
from a pool shared across many Warp users. Registering more identities does
not change the egress IP.
**Consequence: ** "generate 4 distinct Warp identities" is already done and is
NOT the fix. The fix must change * how traffic exits * , not * who the client is * .
## 3. Options
### Option A — Cloudflare Zero Trust dedicated egress IPs (paid, native)
Cloudflare Zero Trust (WARP for Teams) offers **dedicated egress IPs ** as a
paid add-on. With one dedicated egress IP per node and Gateway egress
policies matched on device identity, each node gets a stable, distinct,
organization-owned egress IP.
- Pros: stays on Cloudflare/WireGuard; `netvm-node-up.sh` flow mostly
unchanged (different endpoint + enrollment cert); stable IPs you control;
per-node egress policies give true isolation.
- Cons: paid plan + per-IP add-on (human decision + payment); replaces
wgcf consumer identities with Zero Trust enrollment — the largest config
change of the four options; enrollment certs become new credentials to
guard (same transient-handling rules as Warp keys).
- Human-gated: yes (purchase + org setup).
### Option B — Per-node commercial VPN providers (distinct IPs, non-Warp)
Give each node a WireGuard config from a * different * commercial VPN provider
(e.g. Mullvad, ProtonVPN, IVPN), each with its own exit IP.
- Pros: genuinely distinct egress IPs, guaranteed by construction; no
Cloudflare correlation at all.
- Cons: one account + credentials per provider (5 accounts); non-uniform
configs; recurring cost per provider; abandons the Warp uniformity the
fleet was built on; each provider's ToS / automation posture must be
vetted.
- Human-gated: yes (accounts + payment).
### Option C — Accept shared egress; de-correlate and detect (no new infra)
Keep consumer Warp. Fix the * failure mode * instead of the IP:
1. **Warp egress health stage in `chromebox-watchdog.sh` ** (5th stage after
the existing 4): check `wg show <WG> latest-handshakes` age AND probe
egress from inside the netns
(`ip netns exec warp-<node> curl -m5 https://www.cloudflare.com/cdn-cgi/trace` ).
On stale handshake or failed probe set
`HEALTH_FAIL_REASON="warp egress down"` and take the same relaunch path
as a dead CDP. This converts the invisible partition (CDP green, browser
offline) into a visible, recoverable event.
2. **React-readiness "match" assertion ** in `muse-chat-api.py` before
send/verify: assert title contains `Chat —` and the message-list
container (`role="log"` ) is present and non-skeleton. Fail closed
otherwise. (A partitioned browser otherwise passes naive readiness.)
3. **Per-node rate limiting + jitter ** on muse.ai-bound traffic so the four
nodes don't look like one coordinated client. `bin/rate_limiter.py`
already exists — ensure token buckets are per-node, not shared, and add
jitter to the 5-minute polling loops.
4. **Stagger the fleet polling loops ** (heartbeat, main-loop digests,
health checks) so four nodes never fire in lockstep.
- Pros: no cost, no new credentials, no account changes; directly fixes the
observed failure (invisible partition) and the most likely trigger
(correlated rate-limiting); fully reversible.
- Cons: egress IP stays shared — a hard Cloudflare block of 104.28.195.181
still hits all nodes (but now it is * detected * , * logged distinctly * , and
* recovered from * instead of flapping silently).
- Human-gated: no.
### Option D — Hybrid: per-node egress proxies for sensitive paths only
Keep Warp for general traffic; route muse.ai-bound traffic per node through
distinct SOCKS5/HTTP proxies with distinct IPs (browser `--proxy-server`
per profile).
- Pros: distinct IPs where it matters (muse.ai automation); rest of stack
untouched.
- Cons: proxy accounts + credentials per node; added latency on the hot
path; proxy reliability becomes a new failure domain; Chromium proxy
config per profile adds bookkeeping.
- Human-gated: partially (proxy accounts).
## 4. Recommendation
**Implement Option C now. ** It is the only option that fixes the failure we
actually observed (partitions invisible to the watchdog) with zero new cost
or credentials, and it is fully reversible.
**Keep Option A as the paid path ** if truly distinct, stable egress IPs are
later required (e.g. Cloudflare starts hard-blocking the shared egress IP,
or per-node IP reputation becomes a product need). Option A is the
architecturally clean answer; it just needs a human purchase decision.
**Do not pursue Option B ** unless the fleet leaves Cloudflare for other
reasons — the ops overhead (5 VPN accounts) outweighs the benefit while
Option C mitigates the correlation.
## 5. Identity rotation procedure (for reference)
Not needed for this fix (identities are already distinct), but recorded
because the task asked how identities are generated. Per AGENTS.md, the user
has explicitly authorized operators to run `netvm-new-identity.sh`
("Warp credentials are NOT top secret").
One node at a time, rolling:
``` bash
# on bl, as super (sudo -n via allowlist for netvm-node-up.sh)
NODE = muse # repeat per node
sudo cp /etc/netvm/${ NODE } .conf /root/netvm-backup-$( date +%F) /${ NODE } .conf
sudo rm /etc/netvm/${ NODE } .conf
sudo /home/super/Projects/NetVM/bin/netvm-new-identity.sh ${ NODE }
sudo /home/super/Projects/NetVM/bin/netvm-node-up.sh ${ NODE }
ip netns exec warp-${ NODE } curl -s -m 10 https://api.ipify.org; echo " <- ${ NODE } "
```
`netvm-node-up.sh` is idempotent: it rebuilds the netns, veth pair,
WireGuard interface, and CDP relay from the conf. Wait for the new tunnel's
handshake (`wg show wb-<tag> latest-handshakes` ) before moving to the next
node. Never rotate two nodes simultaneously — keep 4/5 of the fleet up.
## 6. Migration plan (Option C)
Order matters: instrument first, then de-correlate.
1. **Watchdog 5th stage ** — edit `bin/chromebox-watchdog.sh` `healthy()` :
add Warp handshake-age check + in-netns egress probe; set
`HEALTH_FAIL_REASON="warp egress down"` . Deploy to one node (opm) first,
watch one full 5-minute cycle, then roll to all nodes. Commit.
2. **React-readiness assertion ** — add the match check in `muse-chat-api.py`
before send/verify paths. Test against a known-good and a known-bad
(partitioned or landing-page) browser. Commit.
3. **Per-node rate limiter audit ** — confirm `bin/rate_limiter.py` buckets
are keyed per node; add jitter to fleet polling loops (heartbeat timer,
self-main-loop.timer, health checks) so nodes don't fire in lockstep.
Commit.
4. **Stagger verification ** — after deploy, confirm via logs that the four
nodes' 5-minute loops are spread across the minute, not clustered.
Each step is independently committable and revertible. No node downtime is
required for steps 1– 4 (watchdog and API changes take effect on next cycle).
## 7. Migration plan (Option A, if later authorized)
1. Human: purchase Cloudflare Zero Trust plan + 5 dedicated egress IPs.
2. Create the Zero Trust organization; create 5 egress policies, one per
node identity, each pinned to its dedicated egress IP.
3. On bl: enroll each node's Warp client into the org (replaces wgcf
consumer registration); per-node enrollment certs stored like current
confs (root-owned, 0600).
4. Rolling cutover, one node at a time: swap the node's WireGuard config to
the Zero Trust endpoint, run `netvm-node-up.sh <node>` , verify distinct
egress IP (see §8), then proceed.
5. Update NODES.md egress_ip column per node; commit.
## 8. Verification
Per-node, after any change:
``` bash
for n in muse pip 646 opm def; do
ip = $( sudo ip netns exec warp-$n curl -s -m 10 https://api.ipify.org)
echo " $n : $ip "
done | tee /tmp/egress-check-$( date +%F) .log
```
Acceptance for Option C:
- `HEALTH_FAIL_REASON="warp egress down"` appears in watchdog logs when a
tunnel is (test-)partitioned, and the node recovers without manual
intervention.
- Forcing a partition on ONE node (e.g. `wg set <WG> peer <PK>
remove` in a test netns) does not disturb the other nodes' automation.
- Fleet polling loops are measurably staggered in logs.
Acceptance for Option A:
- The 5 egress IPs are pairwise distinct AND stable across 24h
(hourly probe, logged).
- Rate-limiting one node's IP (if testable) leaves the others unaffected.
## 9. Rollback
- **Option C rollback:** revert the watchdog/API commits; the change is
purely additive to health checking — old behavior restored on next
watchdog cycle. No credential or network state to unwind.
- **Identity/config rollback:** ` /root/netvm-backup-<date>/` holds the
pre-change ` /etc/netvm/*.conf` copies; restore and re-run
` netvm-node-up.sh <node>` per node.
- **Option A rollback:** swap each node's WireGuard config back to the
wgcf consumer conf (kept in backup), re-run ` netvm-node-up.sh`.
## 10. Open items (human-gated)
1. Option A requires a paid Cloudflare Zero Trust plan + dedicated egress
IP add-on — purchase decision belongs to the user.
2. If Cloudflare hard-blocks 104.28.195.181 (not just throttles), Option C
detects it but cannot route around it — that event should trigger the
Option A conversation.
3. ` def` node is unassigned; include it in the rollout for uniformity, or
explicitly exclude it and note why.