- Implement muse_hybrid fast gateway in response-harvester.py with ThreadPoolExecutor
- Reduce harvest cycle time from ~1m50s to ~18s; eliminate browser tab-hopping and CDP lock contention
- Add smart active thread filtering in get_monitored_threads to prune dead historical test pipes
- Fix initial watermark ingestion logic so fresh sidechats process first-arrival job responses
- Configure gw_wait window in dm.py (20s on reply:expected) so backend generation is not prematurely severed
- Update box-http-health, box-service-health, box-deep-health to dynamic timestamped sidechats without reuse_key
- Fix heartbeat job to route to dedicated heartbeat channel
- Upgrade exec-constrained.py with subagent.spawn, thread.list, thread.view, pipeline.run ops
- Grant full ops permissions to all fleet agent identities (646, pip, muse, opm)
- Implement bin/box-relay.sh zero-dependency client supporting bearer and SSH signature auth
- Add fast hybrid gateway path to dm.py for sub-2s verified deliveries
- Fix wait=0 handling in super-cli.py subagent deployments
- Add hourly check-in jobs and scheduler for 646, pip, muse
- Document agent tooling and relay APIs in docs/AGENT-TOOLING.md
The unassigned def/dev nodes were being iterated every 5 min by
agent-health.sh (via netvm-registry.py active_nodes()), producing a
perpetual FAIL -> kill -9 -> CRITICAL loop since they can never
"recover". fleet-alert-check.sh iterates the same registry, so this
also stops ghost-node alert noise. The dev registry row itself came
from another session's uncommitted work; this commit keeps their row
and marks it inactive rather than provisioning a node nobody asked
for.
Session: sidechat/chromebox-fixes
- Require 2 CONSECUTIVE muse-chat-api.py API failures before kill -9
(per-node counter in /tmp/agent-health-state, reset on success).
A single 30s API timeout killed 646s healthy browser at 20:48:44 UTC
while its CDP port was still listening.
- Extend post-restart re-check grace to ~60s (15s internal + 45s), matching
chromebox-watchdog.shs proven 60s retry window.
- Add recent_relaunch() guard (mirrors chromebox-watchdog.sh idiom):
skip the kill path when the main browser process launched <2 min ago,
so the two watchdogs can not kill each others fresh browsers.
- check_agent now returns 0/1/2 (healthy/api-fail/port-down); CDP-port
failure still kills immediately. warp-$node checks untouched.
Session: sidechat/chromebox-fixes
Wraps session-pin/unpin/archive/unarchive/rename + threads in the per-agent hybrid gateway transport (netns-isolated, auto-refreshing cookies). Emits the {ok,code,error} contract server.py needs; validates agent/thread/title; never prints secrets, never uses a shell.
Session: sidechat/muse-cli-threads-helper
agent-health.sh used bare node names for 'ip netns exec' but netns are
named warp-<node> since the NetVM layout; the 6189793 CDP-liveness check
always failed ('No such file or directory'), logging false CRITICALs and
kill -9'ing healthy browsers every 5 min. Use warp-$node.
fleet-alert-check.sh: new 5-min critical-condition detector (per-node CDP
liveness via warp-<node> netns). Consecutive-failure state machine:
page after 2 consecutive failures, re-page every 30 min while critical,
quiet-hours-aware (first alert always pages). Emits ALERT/RECOVERY
records to ~/.local/share/fleet-alert/outbox.jsonl for the container
fleet-alert-relay hook; best-effort box-ctl notify to healthy agents.
Session: sidechat/critical-alerting-pipeline
Re-applies three workstreams lost when 793d3d7 committed over uncommitted
edits, reconciled against the parallel track's committed dm.py changes:
- followup-sweeper.py: backfill thread_uuid after successful nudge sends;
record final_nudge_target=main on final-nudge routing (C1/C2)
- response-harvester.py: resolve followups on main-chat replies when
final_nudge_target=main (C3); harvest ALL [RESULT] markers per message
- dm.py: pre-send placement gate (fail closed when post-nav URL lacks the
target thread UUID; skips main) — purely additive over 793d3d7+f268d3d
- sidechat_manager.py: wait_for_chat_list() settle-poll for list population
race (sidebar button renders before titles load)
- new: bin/tests/test_followup_fixes.py (25 tests), bin/placement-audit.py,
bin/dm-log-taxonomy.py, bin/session-probe.py,
docs/SIDECHAT-RELIABILITY.md, docs/UUID-ROTATION.md
Verified: 25/25 tests pass, py_compile clean, sweeper/harvester dry-runs clean.
Known limitation: gate catches wrong-placement, not wrong-mapping (false
autoprovision adopting the parked thread needs a creation check).
- ev(): drain CDP events until response id==1 (was blind recv,
same bug class as box-chat-cdp.py 8d4bfa7 NO_SWITCHER fix)
- cmd_sidechat_use UUID path: use CDP Page.navigate instead of
window.location.href via evaluate; 5 attempts, 15s SPA settle
per attempt (was 3x4s, flaked 1/3 on url_mismatch)
- cmd_sidechat_use name path: retry 3x if not landing on /thread/
Test: 2/2 DM sends succeeded when browser healthy (tests 3-5
hit 646 browser death mid-test, infra issue not nav issue).
Session: sidechat/chromebox-ops
Thread UUIDs rotate -- heartbeat-opm died twice in one day
(5bd5b350 -> 0077e918 -> dead), dropping self_main_loop digests.
SIDCHAT_ALIASES is now intentionally empty; targets fall through to
job-sidechats.json dynamic mappings then name-based sidechat use with
autoprovision, which self-heals. Also removed stale heartbeat-opm and
pipe-9735f2 entries (dead UUID 0077e918) from job-sidechats.json.
Test: dm.py send to heartbeat resolved via name and SENT+VERIFIED
(thread 5f18476d-8994-49e7-a9e0-4838732363fe).
Session: sidechat/chromebox-ops
- Exit 0 on success (was 1 when prompts sent, confusing systemd)
- [!] flag now uses word-boundary regex excluding hyphenated
identities (operator-646 no longer triggers; 'operator needed' does)
- State writes now hold fcntl exclusive lock with fresh reload,
preventing timer check from clobbering enable/disable changes
Session: sidechat/chromebox-ops
- bin/dm.py: remove 646 tasks -> 1e75a740 from SIDCHAT_ALIASES. That UUID is stale (not a valid thread on 646 account; SPA redirects elsewhere, causing misdelivery). Alias now falls through to job-sidechats.json/autoprovision. Do NOT re-add hardcoded UUIDs.
- job-sidechats.json: live autoprovisioned mappings for 646-pip and 646 tasks (2026-10-04).
- bin/box-chat-cdp.py: CDP robustness -- drain events until command response, retry switcher lookup while SPA settles, fresh reconnect per retry (transient NO_SWITCHER on pip/opm 2026-10-04).
Session: sidechat/uuid-stale-fix
Three fixes for flaky main-chat reads (pip/opm):
1. ENSURE: retry chat-switcher lookup 4x with 2s waits instead of
immediate NO_SWITCHER (React may still be rendering).
2. ev(): drain CDP events until matching command id arrives;
previously the first recv() could grab a browser event instead
of our evaluate response, returning None.
3. main(): reconnect fresh websocket on each retry (3 attempts);
reusing a stale ws after page navigation gave dead JS contexts.
Verified: 7/8 reads succeed across all 4 agents; the 1 failure
was THREAD_NOT_FOUND while opm was actively in a side chat,
self-healed on next attempt.
Session: sidechat/chromebox-ops
- self_main_loop.py: 'enabled' map in config (default all True);
do_check skips disabled agents (marks disabled:true in results);
new enable/disable actions with optional --agent.
- box-ctl.py: main-loop enable|disable [--agent <name>] actions,
USAGE updated.
Check skips disabled agents so the loop can be toggled per
Chromebox without stopping the systemd timer.
Session: sidechat/chromebox-ops
Reads each fleet agent's muse.ai Main chat on a 5-min systemd timer
(self-main-loop.timer). On new messages since the per-agent watermark,
posts a concise digest to that agent's prompting sidechat (dm.py, opm
as neutral sender - mirrors box notify), prompting the operator to
check main chat via DM/box. Escapes the main-chat-goes-unread failure
mode. No backfill on first run; silent when nothing new; read/send
failures logged without killing the timer; flock overlap guard.
box-ctl.py: main-loop check | status (fleet section).
Session: sidechat/chromebox-ops
identity-audit-check.sh: proper script replacing the identity-audit-watch
cron inline SSH one-liner. Reads a bl-local cache of the VM audit JSON
(bl cannot SSH to VM; VM hourly audit should push to var/identity-audit.json).
Exits 0 clean, 1 on drift, 2 if cache missing.
box-ctl.py: new identity-audit action returning drift as JSON.
Proper script replacing the inline-SSH latency monitor. Probes each relay /json/version via pinned ports from netvm-names.sh, outputs name:latency_ms:code per node. box-ctl.py cdp-latency runs it and returns JSON.