box fleet status / approvals check misreported every node as STOPPED /
CDP-unreachable from sandboxed shells (own PID+net namespaces: pgrep
blind, no route to 10.201.x.x, no sudo). Fleet was healthy throughout.
- bin/host_evidence.py (new): host watchdog evidence fallback. Recent
timer runs (journal -o json, exact UNIT match) with no newer failure
line in cdp-relay-watchdog.log / chromebox-watchdog.log (both
silent-when-healthy) prove a node is up. def/dev have no watchdog
coverage: browser verdict via chromebox-<node>.log freshness
(alive-only), CDP verdict unknown.
- super-cli.py: effective status/source/evidence per node. Host
evidence decides ONLY the fully-blind pattern (both local probes
negative); live local signals always win. New UNKNOWN badge, [*]
footnote; approvals UNREACHABLE splits into BLIND / OFFLINE(host
agrees) / unreachable-evidence-inconclusive, with honest footer.
proc_alive/cdp_ok keep local-probe meaning; status/source/evidence
are new JSON fields.
- approvals.py: host_cdp_ok flag on the unreachable path.
- agent-health.sh: restart circuit breaker. 3 consecutive futile
restarts (restart leaves agent still failing) opens the circuit:
no more kills for 1800s, ALERT to log+journal, half-open probe
after cooldown, reset on any success. Stops the def murder loop
(57 restarts / 155 API FAILs for an account-layer failure).
- tests/test_fleet_status.py (25), tests/test_agent_health.py (6).
- CHROMEBOX-RUNBOOK.md: blind-shell status + futile-restart sections.
Tests: 98/98 focused green (agent_health + fleet_status +
completion + tool_calls). Live-verified: 4 ACTIVE [*] + 2 UNKNOWN.
Follow-up to 8c39dc3: four scripts hardcoded dev's CDP port instead of
reading netvm-registry.py. Aligned with the corrected registry (9455).
Session: sidechat/dev-cdp-partition
- Upgrade exec-constrained.py with subagent.spawn, thread.list, thread.view, pipeline.run ops
- Grant full ops permissions to all fleet agent identities (646, pip, muse, opm)
- Implement bin/box-relay.sh zero-dependency client supporting bearer and SSH signature auth
- Add fast hybrid gateway path to dm.py for sub-2s verified deliveries
- Fix wait=0 handling in super-cli.py subagent deployments
- Add hourly check-in jobs and scheduler for 646, pip, muse
- Document agent tooling and relay APIs in docs/AGENT-TOOLING.md
- Add stop_pipeline and prune_pipelines to bin/pipeline_engine.py
- Support prefix matching and custom cancellation reason
- Wire 'super pipeline stop <run_id>' and 'super pipeline prune [--max-age H]' into bin/super-cli.py
- Added CHAT_POLICY.md and README.md banner enforcing sidechat-first and file-transfer-first rules.
- Added strict Main Chat block to super dm send and super dm wo with --allow-main-chat override.
- Implemented file transfer staging and metadata registry in super dm send-file and super dm files (with clean subcommand).
- Added full job lifecycle management (show, create, enable, disable, delete, run --follow) to super-cli.py and box-ctl.py.
- Audited all jobs in jobs/*.json and redirected automated dispatches away from Main Chat.
- Hardened chromebox-watchdog.sh with systemd user session environment exports and stale singleton cleanup.
- Added compose_check command choice to muse-chat-api.py.