feat: setup-fed watchdog supervision for all registry nodes

Close the def/dev supervision gap at the source: every node brought
up gets watched, and every supervisor enumerates the registry.

- bin/ensure-node-supervision.sh (new, idempotent): appends the
  NODES.md row (netvm-names port, honors CDP_PORT_OVERRIDE so it
  never fights provision's picker) and installs/enables
  chromebox-watchdog-<node>.timer. --all heals drift (registry +
  /etc/netvm identities). Template verified byte-identical to the
  installed def unit.
- netvm-node-up.sh: calls ensure (non-fatal) at the end. Provision
  and the onboarding pipeline reach it transitively.
- relay-health-check.sh, cdp-latency-check.sh: registry-driven
  watched_nodes() + LIB_ONLY guards (were hardcoded 4 nodes).
- tests/test_node_supervision.py (6): row add/idempotent/override,
  timer render, node-up wiring, both watched_nodes().
- CHROMEBOX-RUNBOOK.md: setup-fed supervision section.

Pairs with the registry-driven relay/chromebox watchdogs: new rows
are picked up on the next run with no per-node code edits.
This commit is contained in:
Muse Sidechat
2026-10-06 19:28:29 +00:00
parent c9143a558b
commit 66c8900a58
6 changed files with 249 additions and 3 deletions
+5
View File
@@ -104,3 +104,8 @@ else
EGRESS=$(nsexec curl -sk --max-time 15 'https://1.1.1.1/cdn-cgi/trace' 2>/dev/null | grep -oP '^ip=\K.*' || true)
fi
echo "node=$NODE netns=$NETNS ifaces=$WG/$VETH egress=${EGRESS:-unknown}"
# Feed the watchdogs: registry row + chromebox timer (non-fatal — the
# node is up regardless, and supervision heals on the next run).
"$SCRIPT_DIR/ensure-node-supervision.sh" "$NODE" \
|| echo "supervision ensure failed for $NODE (non-fatal)" >&2