The unassigned def/dev nodes were being iterated every 5 min by
agent-health.sh (via netvm-registry.py active_nodes()), producing a
perpetual FAIL -> kill -9 -> CRITICAL loop since they can never
"recover". fleet-alert-check.sh iterates the same registry, so this
also stops ghost-node alert noise. The dev registry row itself came
from another session's uncommitted work; this commit keeps their row
and marks it inactive rather than provisioning a node nobody asked
for.
Session: sidechat/chromebox-fixes
New bin/netvm-registry.py parses NODES.md (node -> cdp_port); every
consumer reads from it instead of hardcoding:
- onboard-driver.py: CDP_PORTS dict -> registry lookup (new nodes work)
- muse-signin.py: hardcoded 9410 -> --node/--cdp-port args
- muse-chat-api.py: hardcoded ACCOUNTS -> registry-built
- agent-health.sh: hardcoded muse/pip blocks -> loop over all active
nodes (646 and opm now get health coverage too)
NODES.md: record opm node (9440).