Close the def/dev supervision gap at the source: every node brought
up gets watched, and every supervisor enumerates the registry.
- bin/ensure-node-supervision.sh (new, idempotent): appends the
NODES.md row (netvm-names port, honors CDP_PORT_OVERRIDE so it
never fights provision's picker) and installs/enables
chromebox-watchdog-<node>.timer. --all heals drift (registry +
/etc/netvm identities). Template verified byte-identical to the
installed def unit.
- netvm-node-up.sh: calls ensure (non-fatal) at the end. Provision
and the onboarding pipeline reach it transitively.
- relay-health-check.sh, cdp-latency-check.sh: registry-driven
watched_nodes() + LIB_ONLY guards (were hardcoded 4 nodes).
- tests/test_node_supervision.py (6): row add/idempotent/override,
timer render, node-up wiring, both watched_nodes().
- CHROMEBOX-RUNBOOK.md: setup-fed supervision section.
Pairs with the registry-driven relay/chromebox watchdogs: new rows
are picked up on the next run with no per-node code edits.
box fleet status / approvals check misreported every node as STOPPED /
CDP-unreachable from sandboxed shells (own PID+net namespaces: pgrep
blind, no route to 10.201.x.x, no sudo). Fleet was healthy throughout.
- bin/host_evidence.py (new): host watchdog evidence fallback. Recent
timer runs (journal -o json, exact UNIT match) with no newer failure
line in cdp-relay-watchdog.log / chromebox-watchdog.log (both
silent-when-healthy) prove a node is up. def/dev have no watchdog
coverage: browser verdict via chromebox-<node>.log freshness
(alive-only), CDP verdict unknown.
- super-cli.py: effective status/source/evidence per node. Host
evidence decides ONLY the fully-blind pattern (both local probes
negative); live local signals always win. New UNKNOWN badge, [*]
footnote; approvals UNREACHABLE splits into BLIND / OFFLINE(host
agrees) / unreachable-evidence-inconclusive, with honest footer.
proc_alive/cdp_ok keep local-probe meaning; status/source/evidence
are new JSON fields.
- approvals.py: host_cdp_ok flag on the unreachable path.
- agent-health.sh: restart circuit breaker. 3 consecutive futile
restarts (restart leaves agent still failing) opens the circuit:
no more kills for 1800s, ALERT to log+journal, half-open probe
after cooldown, reset on any success. Stops the def murder loop
(57 restarts / 155 API FAILs for an account-layer failure).
- tests/test_fleet_status.py (25), tests/test_agent_health.py (6).
- CHROMEBOX-RUNBOOK.md: blind-shell status + futile-restart sections.
Tests: 98/98 focused green (agent_health + fleet_status +
completion + tool_calls). Live-verified: 4 ACTIVE [*] + 2 UNKNOWN.
echo "$list" | grep -q under set -o pipefail exits 141 whenever grep
matches before echo finishes writing, so healthy browsers were reported
'CDP up but no page target' and killed every 2 min fleet-wide (load 19+).
Use [[ == *glob* ]] (no pipe, no race) for the page/muse.ai stages, and
grep -c (reads to EOF, never early-exits) for the netns check.
The 2026-10-05 fix note claimed dev/def were removed from the pool but the
code still listed [dev, def, muse]. The harvester's every-minute systemd
dispatch raced the pool daemon claiming slots as dev/def, which refuse on
attribution grounds (dev FAILs fast, def freezes until the 60-min reaper) -
the fleet-wide dev-FAIL-fast + def-frozen pattern on every swarm.
Pool now matches swarm_worker/daemon.py WORKER_POOL exactly: muse only,
the single fully authenticated auxiliary worker. Reporting/harvest paths
untouched.
- daemon: route non-shell slot tasks to a fresh subagent session on dev/def/muse
via muse_hybrid; add bin/ to sys.path so muse_hybrid/prompt_envelope import
under the tmux supervisor
- prompt_envelope: add wrap_subagent_task() - plain task assignment without
[TOOL tmux/swarm/cron] meta tags (subagents refused those as relayed test traffic)
- box-ctl/reporter: swarm-attach accepts optional session-id, stored as
slot.subagent_session_id
- executor: _looks_like_shell requires an existing executable (prose like
'verify ...' no longer misread as shell)
- poller: pick up pending swarms as well as running
- response-harvester: monitor running slot subagent sessions from swarms.json,
allow '/' in RESULT/VERB job ids, archive ephemeral threads on any verdict
(OK or FAIL), reap slot sessions for terminal swarms
Follow-up to 8c39dc3: four scripts hardcoded dev's CDP port instead of
reading netvm-registry.py. Aligned with the corrected registry (9455).
Session: sidechat/dev-cdp-partition
NODES.md advertised dev's CDP on 9460, but the fleet-facing path is the
relay netvm-cdp-relay.py 10.201.36.2:9455 -> 127.0.0.1:9455, and
netvm-names.sh never pinned dev (hash default = 9455). The watchdog kept
relaunching dev's browser on 9460 while the relay sat orphaned on 9455.
- NODES.md: dev cdp_port 9460 -> 9455 (registry is the watchdog's source)
- bin/netvm-names.sh: pin def=9450, dev=9455 (was hash-default only)
Reported by 646 as xop-e09 partition; verified hop-by-hop before fix.
Session: sidechat/dev-cdp-partition