service.status runs bl-local (.timer -> user bus, .service -> system bus).
board.service/caddy.service are inactive on bl by design -- they run on the
VM and are covered by box-health-check.sh over the SSH chain.
response-harvester.timer lives on the bl user bus, not the VM.
- Implement muse_hybrid fast gateway in response-harvester.py with ThreadPoolExecutor
- Reduce harvest cycle time from ~1m50s to ~18s; eliminate browser tab-hopping and CDP lock contention
- Add smart active thread filtering in get_monitored_threads to prune dead historical test pipes
- Fix initial watermark ingestion logic so fresh sidechats process first-arrival job responses
- Configure gw_wait window in dm.py (20s on reply:expected) so backend generation is not prematurely severed
- Update box-http-health, box-service-health, box-deep-health to dynamic timestamped sidechats without reuse_key
- Fix heartbeat job to route to dedicated heartbeat channel
- Upgrade exec-constrained.py with subagent.spawn, thread.list, thread.view, pipeline.run ops
- Grant full ops permissions to all fleet agent identities (646, pip, muse, opm)
- Implement bin/box-relay.sh zero-dependency client supporting bearer and SSH signature auth
- Add fast hybrid gateway path to dm.py for sub-2s verified deliveries
- Fix wait=0 handling in super-cli.py subagent deployments
- Add hourly check-in jobs and scheduler for 646, pip, muse
- Document agent tooling and relay APIs in docs/AGENT-TOOLING.md
The unassigned def/dev nodes were being iterated every 5 min by
agent-health.sh (via netvm-registry.py active_nodes()), producing a
perpetual FAIL -> kill -9 -> CRITICAL loop since they can never
"recover". fleet-alert-check.sh iterates the same registry, so this
also stops ghost-node alert noise. The dev registry row itself came
from another session's uncommitted work; this commit keeps their row
and marks it inactive rather than provisioning a node nobody asked
for.
Session: sidechat/chromebox-fixes
- Require 2 CONSECUTIVE muse-chat-api.py API failures before kill -9
(per-node counter in /tmp/agent-health-state, reset on success).
A single 30s API timeout killed 646s healthy browser at 20:48:44 UTC
while its CDP port was still listening.
- Extend post-restart re-check grace to ~60s (15s internal + 45s), matching
chromebox-watchdog.shs proven 60s retry window.
- Add recent_relaunch() guard (mirrors chromebox-watchdog.sh idiom):
skip the kill path when the main browser process launched <2 min ago,
so the two watchdogs can not kill each others fresh browsers.
- check_agent now returns 0/1/2 (healthy/api-fail/port-down); CDP-port
failure still kills immediately. warp-$node checks untouched.
Session: sidechat/chromebox-fixes