Close the def/dev supervision gap at the source: every node brought
up gets watched, and every supervisor enumerates the registry.
- bin/ensure-node-supervision.sh (new, idempotent): appends the
NODES.md row (netvm-names port, honors CDP_PORT_OVERRIDE so it
never fights provision's picker) and installs/enables
chromebox-watchdog-<node>.timer. --all heals drift (registry +
/etc/netvm identities). Template verified byte-identical to the
installed def unit.
- netvm-node-up.sh: calls ensure (non-fatal) at the end. Provision
and the onboarding pipeline reach it transitively.
- relay-health-check.sh, cdp-latency-check.sh: registry-driven
watched_nodes() + LIB_ONLY guards (were hardcoded 4 nodes).
- tests/test_node_supervision.py (6): row add/idempotent/override,
timer render, node-up wiring, both watched_nodes().
- CHROMEBOX-RUNBOOK.md: setup-fed supervision section.
Pairs with the registry-driven relay/chromebox watchdogs: new rows
are picked up on the next run with no per-node code edits.
box fleet status / approvals check misreported every node as STOPPED /
CDP-unreachable from sandboxed shells (own PID+net namespaces: pgrep
blind, no route to 10.201.x.x, no sudo). Fleet was healthy throughout.
- bin/host_evidence.py (new): host watchdog evidence fallback. Recent
timer runs (journal -o json, exact UNIT match) with no newer failure
line in cdp-relay-watchdog.log / chromebox-watchdog.log (both
silent-when-healthy) prove a node is up. def/dev have no watchdog
coverage: browser verdict via chromebox-<node>.log freshness
(alive-only), CDP verdict unknown.
- super-cli.py: effective status/source/evidence per node. Host
evidence decides ONLY the fully-blind pattern (both local probes
negative); live local signals always win. New UNKNOWN badge, [*]
footnote; approvals UNREACHABLE splits into BLIND / OFFLINE(host
agrees) / unreachable-evidence-inconclusive, with honest footer.
proc_alive/cdp_ok keep local-probe meaning; status/source/evidence
are new JSON fields.
- approvals.py: host_cdp_ok flag on the unreachable path.
- agent-health.sh: restart circuit breaker. 3 consecutive futile
restarts (restart leaves agent still failing) opens the circuit:
no more kills for 1800s, ALERT to log+journal, half-open probe
after cooldown, reset on any success. Stops the def murder loop
(57 restarts / 155 API FAILs for an account-layer failure).
- tests/test_fleet_status.py (25), tests/test_agent_health.py (6).
- CHROMEBOX-RUNBOOK.md: blind-shell status + futile-restart sections.
Tests: 98/98 focused green (agent_health + fleet_status +
completion + tool_calls). Live-verified: 4 ACTIVE [*] + 2 UNKNOWN.