feat: setup-fed watchdog supervision for all registry nodes
Close the def/dev supervision gap at the source: every node brought up gets watched, and every supervisor enumerates the registry. - bin/ensure-node-supervision.sh (new, idempotent): appends the NODES.md row (netvm-names port, honors CDP_PORT_OVERRIDE so it never fights provision's picker) and installs/enables chromebox-watchdog-<node>.timer. --all heals drift (registry + /etc/netvm identities). Template verified byte-identical to the installed def unit. - netvm-node-up.sh: calls ensure (non-fatal) at the end. Provision and the onboarding pipeline reach it transitively. - relay-health-check.sh, cdp-latency-check.sh: registry-driven watched_nodes() + LIB_ONLY guards (were hardcoded 4 nodes). - tests/test_node_supervision.py (6): row add/idempotent/override, timer render, node-up wiring, both watched_nodes(). - CHROMEBOX-RUNBOOK.md: setup-fed supervision section. Pairs with the registry-driven relay/chromebox watchdogs: new rows are picked up on the next run with no per-node code edits.
This commit is contained in:
@@ -104,3 +104,8 @@ else
|
||||
EGRESS=$(nsexec curl -sk --max-time 15 'https://1.1.1.1/cdn-cgi/trace' 2>/dev/null | grep -oP '^ip=\K.*' || true)
|
||||
fi
|
||||
echo "node=$NODE netns=$NETNS ifaces=$WG/$VETH egress=${EGRESS:-unknown}"
|
||||
|
||||
# Feed the watchdogs: registry row + chromebox timer (non-fatal — the
|
||||
# node is up regardless, and supervision heals on the next run).
|
||||
"$SCRIPT_DIR/ensure-node-supervision.sh" "$NODE" \
|
||||
|| echo "supervision ensure failed for $NODE (non-fatal)" >&2
|
||||
|
||||
Reference in New Issue
Block a user