- Upgrade exec-constrained.py with subagent.spawn, thread.list, thread.view, pipeline.run ops
- Grant full ops permissions to all fleet agent identities (646, pip, muse, opm)
- Implement bin/box-relay.sh zero-dependency client supporting bearer and SSH signature auth
- Add fast hybrid gateway path to dm.py for sub-2s verified deliveries
- Fix wait=0 handling in super-cli.py subagent deployments
- Add hourly check-in jobs and scheduler for 646, pip, muse
- Document agent tooling and relay APIs in docs/AGENT-TOOLING.md
- Require 2 CONSECUTIVE muse-chat-api.py API failures before kill -9
(per-node counter in /tmp/agent-health-state, reset on success).
A single 30s API timeout killed 646s healthy browser at 20:48:44 UTC
while its CDP port was still listening.
- Extend post-restart re-check grace to ~60s (15s internal + 45s), matching
chromebox-watchdog.shs proven 60s retry window.
- Add recent_relaunch() guard (mirrors chromebox-watchdog.sh idiom):
skip the kill path when the main browser process launched <2 min ago,
so the two watchdogs can not kill each others fresh browsers.
- check_agent now returns 0/1/2 (healthy/api-fail/port-down); CDP-port
failure still kills immediately. warp-$node checks untouched.
Session: sidechat/chromebox-fixes
Wraps session-pin/unpin/archive/unarchive/rename + threads in the per-agent hybrid gateway transport (netns-isolated, auto-refreshing cookies). Emits the {ok,code,error} contract server.py needs; validates agent/thread/title; never prints secrets, never uses a shell.
Session: sidechat/muse-cli-threads-helper
agent-health.sh used bare node names for 'ip netns exec' but netns are
named warp-<node> since the NetVM layout; the 6189793 CDP-liveness check
always failed ('No such file or directory'), logging false CRITICALs and
kill -9'ing healthy browsers every 5 min. Use warp-$node.
fleet-alert-check.sh: new 5-min critical-condition detector (per-node CDP
liveness via warp-<node> netns). Consecutive-failure state machine:
page after 2 consecutive failures, re-page every 30 min while critical,
quiet-hours-aware (first alert always pages). Emits ALERT/RECOVERY
records to ~/.local/share/fleet-alert/outbox.jsonl for the container
fleet-alert-relay hook; best-effort box-ctl notify to healthy agents.
Session: sidechat/critical-alerting-pipeline
Re-applies three workstreams lost when 793d3d7 committed over uncommitted
edits, reconciled against the parallel track's committed dm.py changes:
- followup-sweeper.py: backfill thread_uuid after successful nudge sends;
record final_nudge_target=main on final-nudge routing (C1/C2)
- response-harvester.py: resolve followups on main-chat replies when
final_nudge_target=main (C3); harvest ALL [RESULT] markers per message
- dm.py: pre-send placement gate (fail closed when post-nav URL lacks the
target thread UUID; skips main) — purely additive over 793d3d7+f268d3d
- sidechat_manager.py: wait_for_chat_list() settle-poll for list population
race (sidebar button renders before titles load)
- new: bin/tests/test_followup_fixes.py (25 tests), bin/placement-audit.py,
bin/dm-log-taxonomy.py, bin/session-probe.py,
docs/SIDECHAT-RELIABILITY.md, docs/UUID-ROTATION.md
Verified: 25/25 tests pass, py_compile clean, sweeper/harvester dry-runs clean.
Known limitation: gate catches wrong-placement, not wrong-mapping (false
autoprovision adopting the parked thread needs a creation check).
- ev(): drain CDP events until response id==1 (was blind recv,
same bug class as box-chat-cdp.py 8d4bfa7 NO_SWITCHER fix)
- cmd_sidechat_use UUID path: use CDP Page.navigate instead of
window.location.href via evaluate; 5 attempts, 15s SPA settle
per attempt (was 3x4s, flaked 1/3 on url_mismatch)
- cmd_sidechat_use name path: retry 3x if not landing on /thread/
Test: 2/2 DM sends succeeded when browser healthy (tests 3-5
hit 646 browser death mid-test, infra issue not nav issue).
Session: sidechat/chromebox-ops
Thread UUIDs rotate -- heartbeat-opm died twice in one day
(5bd5b350 -> 0077e918 -> dead), dropping self_main_loop digests.
SIDCHAT_ALIASES is now intentionally empty; targets fall through to
job-sidechats.json dynamic mappings then name-based sidechat use with
autoprovision, which self-heals. Also removed stale heartbeat-opm and
pipe-9735f2 entries (dead UUID 0077e918) from job-sidechats.json.
Test: dm.py send to heartbeat resolved via name and SENT+VERIFIED
(thread 5f18476d-8994-49e7-a9e0-4838732363fe).
Session: sidechat/chromebox-ops
- Exit 0 on success (was 1 when prompts sent, confusing systemd)
- [!] flag now uses word-boundary regex excluding hyphenated
identities (operator-646 no longer triggers; 'operator needed' does)
- State writes now hold fcntl exclusive lock with fresh reload,
preventing timer check from clobbering enable/disable changes
Session: sidechat/chromebox-ops
- bin/dm.py: remove 646 tasks -> 1e75a740 from SIDCHAT_ALIASES. That UUID is stale (not a valid thread on 646 account; SPA redirects elsewhere, causing misdelivery). Alias now falls through to job-sidechats.json/autoprovision. Do NOT re-add hardcoded UUIDs.
- job-sidechats.json: live autoprovisioned mappings for 646-pip and 646 tasks (2026-10-04).
- bin/box-chat-cdp.py: CDP robustness -- drain events until command response, retry switcher lookup while SPA settles, fresh reconnect per retry (transient NO_SWITCHER on pip/opm 2026-10-04).
Session: sidechat/uuid-stale-fix
Three fixes for flaky main-chat reads (pip/opm):
1. ENSURE: retry chat-switcher lookup 4x with 2s waits instead of
immediate NO_SWITCHER (React may still be rendering).
2. ev(): drain CDP events until matching command id arrives;
previously the first recv() could grab a browser event instead
of our evaluate response, returning None.
3. main(): reconnect fresh websocket on each retry (3 attempts);
reusing a stale ws after page navigation gave dead JS contexts.
Verified: 7/8 reads succeed across all 4 agents; the 1 failure
was THREAD_NOT_FOUND while opm was actively in a side chat,
self-healed on next attempt.
Session: sidechat/chromebox-ops
- self_main_loop.py: 'enabled' map in config (default all True);
do_check skips disabled agents (marks disabled:true in results);
new enable/disable actions with optional --agent.
- box-ctl.py: main-loop enable|disable [--agent <name>] actions,
USAGE updated.
Check skips disabled agents so the loop can be toggled per
Chromebox without stopping the systemd timer.
Session: sidechat/chromebox-ops
Reads each fleet agent's muse.ai Main chat on a 5-min systemd timer
(self-main-loop.timer). On new messages since the per-agent watermark,
posts a concise digest to that agent's prompting sidechat (dm.py, opm
as neutral sender - mirrors box notify), prompting the operator to
check main chat via DM/box. Escapes the main-chat-goes-unread failure
mode. No backfill on first run; silent when nothing new; read/send
failures logged without killing the timer; flock overlap guard.
box-ctl.py: main-loop check | status (fleet section).
Session: sidechat/chromebox-ops
identity-audit-check.sh: proper script replacing the identity-audit-watch
cron inline SSH one-liner. Reads a bl-local cache of the VM audit JSON
(bl cannot SSH to VM; VM hourly audit should push to var/identity-audit.json).
Exits 0 clean, 1 on drift, 2 if cache missing.
box-ctl.py: new identity-audit action returning drift as JSON.
Proper script replacing the inline-SSH latency monitor. Probes each relay /json/version via pinned ports from netvm-names.sh, outputs name:latency_ms:code per node. box-ctl.py cdp-latency runs it and returns JSON.
Self-contained FAILED-relaunch scanner for chromebox-watchdog.log
with watermark at NETVM_ROOT/watchdog-alert-watermark.txt.
Exits 0 quiet when clean, 1 with new lines printed when alerting.
The box-ctl.py watchdog-alerts action (in 2db0d92) drives this script.
Session: sidechat/chromebox-ops
Converts the cdp-relay-health-monitor from inline SSH (nested quoting
bugs) to a proper self-contained script on bl. Uses pinned ports from
netvm-names.sh as single source of truth.
New files/actions:
- bin/relay-health-check.sh: checks all four CDP relays, exits 0/1
- box-ctl.py relay-health: JSON wrapper for the Box CLI
Session: sidechat/chromebox-ops
restart_relay() ran pkill without sudo, but relay processes are
root-owned and the timer runs as User=super. The kill failed EPERM
(silently swallowed by || true), the old relay kept running, and
SO_REUSEADDR let the replacement double-bind the same port. The
post-restart health check then passed and logged "restarted OK"
when nothing was actually restarted.
Adding sudo -n to the pkill, matching the sudo -n ip netns exec
already used to start the relay.
Session: sidechat/chromebox-ops
The relaunch-loop guard only protected the kill step, not the relaunch.
When CDP was unreachable on a slow-starting browser, the watchdog would
invoke netvm-chrome.sh (which kills the existing browser) before checking
if it was recently launched — piling up 5 chromiums on opm.
Now the <120s check runs before the relaunch and skips the entire cycle.
Also fixed the PID check to match only the main browser process
(--remote-debugging-port, excluding --type= renderer/gpu children).
Session: sidechat/chromebox-ops
netvm-node-up.sh was starting the CDP relay with the hash-derived
CDP_PORT while cdp-relay-watchdog.sh pinned muse->9410, pip->9420,
646->9430, opm->9440. Since netvm-chrome.sh calls netvm-node-up.sh
on every browser relaunch, each restart spawned a wrong-port zombie
relay (9353, 10355, 10239...).
The pinning now lives in netvm_names() itself, making it the single
source of truth for all consumers (node-up, chrome, cdp, accounts,
watchdog). CDP_PORT_OVERRIDE still takes precedence for future nodes.
Session: sidechat/chromebox-ops
Identifier, OTP code, and account-name values replaced with [redacted] in all progress prints. Line structure and markers preserved (APPROVAL_NEEDED/NEEDS_HUMAN/SUCCESS prefixes intact); exit codes 0/2/3/4 unchanged, no consumer parses stdout. Defense in depth: onboard-driver already scrubs its passthrough; this covers manual operator flows too.
- Redact identifier/code (>=4 chars) from muse-signin.py stdout/stderr before passthrough (it echoes OTP input).
- Catch TimeoutExpired: its str() includes argv with --email/--otp; print generic timeout instead.
- Log email_masked instead of raw email to job-log.jsonl (rc=4 branch).
dm_send already knows the thread UUID at send time (alias resolution, nav URL capture, autoprovision capture) but dropped it: the only stdout channel the board parses carried no thread info, so every auto-logged DM row had thread_uuid null.
Now prints '... SENT and VERIFIED thread=<uuid>' when known, unchanged line otherwise (main-chat sends, UUID-capture failures stay honestly null -- never fabricated). Backward compatible: board _BOX_DM_SENT_RE has no end anchor; bl consumers (job-dispatch.py, siphon-bl.py) use substring matching.
Worktree note: unrelated uncommitted changes remain (job_id followup-correlation hunks in this file, followup-sweeper.py, response-harvester.py, untracked helpers) -- not mine, not staged.
Session: sidechat/box-dms-ui
- Add stop_pipeline and prune_pipelines to bin/pipeline_engine.py
- Support prefix matching and custom cancellation reason
- Wire 'super pipeline stop <run_id>' and 'super pipeline prune [--max-age H]' into bin/super-cli.py
- Add bin/pipeline_engine.py for persistent multi-agent execution tracking in pipelines.json
- Add jobs/pipe-demo-step1.json and jobs/pipe-demo-step2.json demo pipeline definitions
- Use ev1 in bin/muse-chat-api.py across send/messages/compose/create to avoid dropping return values on CDP event chatter
- Support Muse unconfirmed signup error handling in bin/muse-signin.py
- Add runtime state and telemetry files to .gitignore
- Track dynamic pipe sidechat mappings in job-sidechats.json
- Add bin/gravity.py loop diagnostics, reconstruction, and progressive remediation
- Wire hard-break alerting to job-log audit and operator direct message
- Add comprehensive architecture and operational specification in docs/LOOP-MANAGEMENT.md
- Add sidechat thread auto-provisioning fallback on 'Navigated to: None' in bin/dm.py
- Support Muse unconfirmed signup error handling in bin/muse-signin.py
- Track dynamic pipe sidechat mappings in job-sidechats.json
- Added CHAT_POLICY.md and README.md banner enforcing sidechat-first and file-transfer-first rules.
- Added strict Main Chat block to super dm send and super dm wo with --allow-main-chat override.
- Implemented file transfer staging and metadata registry in super dm send-file and super dm files (with clean subcommand).
- Added full job lifecycle management (show, create, enable, disable, delete, run --follow) to super-cli.py and box-ctl.py.
- Audited all jobs in jobs/*.json and redirected automated dispatches away from Main Chat.
- Hardened chromebox-watchdog.sh with systemd user session environment exports and stale singleton cleanup.
- Added compose_check command choice to muse-chat-api.py.
Two fixes: (1) retry health check 4x with 15s gaps after relaunch instead of single 25s check; (2) skip kill if browser launched <2min ago (probably still starting). Prevents watchdog from killing a working-but-slow browser.
Session: sidechat/chromebox-ops