Three fixes for flaky main-chat reads (pip/opm):
1. ENSURE: retry chat-switcher lookup 4x with 2s waits instead of
immediate NO_SWITCHER (React may still be rendering).
2. ev(): drain CDP events until matching command id arrives;
previously the first recv() could grab a browser event instead
of our evaluate response, returning None.
3. main(): reconnect fresh websocket on each retry (3 attempts);
reusing a stale ws after page navigation gave dead JS contexts.
Verified: 7/8 reads succeed across all 4 agents; the 1 failure
was THREAD_NOT_FOUND while opm was actively in a side chat,
self-healed on next attempt.
Session: sidechat/chromebox-ops
- self_main_loop.py: 'enabled' map in config (default all True);
do_check skips disabled agents (marks disabled:true in results);
new enable/disable actions with optional --agent.
- box-ctl.py: main-loop enable|disable [--agent <name>] actions,
USAGE updated.
Check skips disabled agents so the loop can be toggled per
Chromebox without stopping the systemd timer.
Session: sidechat/chromebox-ops
Reads each fleet agent's muse.ai Main chat on a 5-min systemd timer
(self-main-loop.timer). On new messages since the per-agent watermark,
posts a concise digest to that agent's prompting sidechat (dm.py, opm
as neutral sender - mirrors box notify), prompting the operator to
check main chat via DM/box. Escapes the main-chat-goes-unread failure
mode. No backfill on first run; silent when nothing new; read/send
failures logged without killing the timer; flock overlap guard.
box-ctl.py: main-loop check | status (fleet section).
Session: sidechat/chromebox-ops
identity-audit-check.sh: proper script replacing the identity-audit-watch
cron inline SSH one-liner. Reads a bl-local cache of the VM audit JSON
(bl cannot SSH to VM; VM hourly audit should push to var/identity-audit.json).
Exits 0 clean, 1 on drift, 2 if cache missing.
box-ctl.py: new identity-audit action returning drift as JSON.
Proper script replacing the inline-SSH latency monitor. Probes each relay /json/version via pinned ports from netvm-names.sh, outputs name:latency_ms:code per node. box-ctl.py cdp-latency runs it and returns JSON.
Self-contained FAILED-relaunch scanner for chromebox-watchdog.log
with watermark at NETVM_ROOT/watchdog-alert-watermark.txt.
Exits 0 quiet when clean, 1 with new lines printed when alerting.
The box-ctl.py watchdog-alerts action (in 2db0d92) drives this script.
Session: sidechat/chromebox-ops
Converts the cdp-relay-health-monitor from inline SSH (nested quoting
bugs) to a proper self-contained script on bl. Uses pinned ports from
netvm-names.sh as single source of truth.
New files/actions:
- bin/relay-health-check.sh: checks all four CDP relays, exits 0/1
- box-ctl.py relay-health: JSON wrapper for the Box CLI
Session: sidechat/chromebox-ops
restart_relay() ran pkill without sudo, but relay processes are
root-owned and the timer runs as User=super. The kill failed EPERM
(silently swallowed by || true), the old relay kept running, and
SO_REUSEADDR let the replacement double-bind the same port. The
post-restart health check then passed and logged "restarted OK"
when nothing was actually restarted.
Adding sudo -n to the pkill, matching the sudo -n ip netns exec
already used to start the relay.
Session: sidechat/chromebox-ops
The relaunch-loop guard only protected the kill step, not the relaunch.
When CDP was unreachable on a slow-starting browser, the watchdog would
invoke netvm-chrome.sh (which kills the existing browser) before checking
if it was recently launched — piling up 5 chromiums on opm.
Now the <120s check runs before the relaunch and skips the entire cycle.
Also fixed the PID check to match only the main browser process
(--remote-debugging-port, excluding --type= renderer/gpu children).
Session: sidechat/chromebox-ops
netvm-node-up.sh was starting the CDP relay with the hash-derived
CDP_PORT while cdp-relay-watchdog.sh pinned muse->9410, pip->9420,
646->9430, opm->9440. Since netvm-chrome.sh calls netvm-node-up.sh
on every browser relaunch, each restart spawned a wrong-port zombie
relay (9353, 10355, 10239...).
The pinning now lives in netvm_names() itself, making it the single
source of truth for all consumers (node-up, chrome, cdp, accounts,
watchdog). CDP_PORT_OVERRIDE still takes precedence for future nodes.
Session: sidechat/chromebox-ops
Identifier, OTP code, and account-name values replaced with [redacted] in all progress prints. Line structure and markers preserved (APPROVAL_NEEDED/NEEDS_HUMAN/SUCCESS prefixes intact); exit codes 0/2/3/4 unchanged, no consumer parses stdout. Defense in depth: onboard-driver already scrubs its passthrough; this covers manual operator flows too.
- Redact identifier/code (>=4 chars) from muse-signin.py stdout/stderr before passthrough (it echoes OTP input).
- Catch TimeoutExpired: its str() includes argv with --email/--otp; print generic timeout instead.
- Log email_masked instead of raw email to job-log.jsonl (rc=4 branch).
dm_send already knows the thread UUID at send time (alias resolution, nav URL capture, autoprovision capture) but dropped it: the only stdout channel the board parses carried no thread info, so every auto-logged DM row had thread_uuid null.
Now prints '... SENT and VERIFIED thread=<uuid>' when known, unchanged line otherwise (main-chat sends, UUID-capture failures stay honestly null -- never fabricated). Backward compatible: board _BOX_DM_SENT_RE has no end anchor; bl consumers (job-dispatch.py, siphon-bl.py) use substring matching.
Worktree note: unrelated uncommitted changes remain (job_id followup-correlation hunks in this file, followup-sweeper.py, response-harvester.py, untracked helpers) -- not mine, not staged.
Session: sidechat/box-dms-ui
- Add stop_pipeline and prune_pipelines to bin/pipeline_engine.py
- Support prefix matching and custom cancellation reason
- Wire 'super pipeline stop <run_id>' and 'super pipeline prune [--max-age H]' into bin/super-cli.py
- Add bin/pipeline_engine.py for persistent multi-agent execution tracking in pipelines.json
- Add jobs/pipe-demo-step1.json and jobs/pipe-demo-step2.json demo pipeline definitions
- Use ev1 in bin/muse-chat-api.py across send/messages/compose/create to avoid dropping return values on CDP event chatter
- Support Muse unconfirmed signup error handling in bin/muse-signin.py
- Add runtime state and telemetry files to .gitignore
- Track dynamic pipe sidechat mappings in job-sidechats.json
- Add bin/gravity.py loop diagnostics, reconstruction, and progressive remediation
- Wire hard-break alerting to job-log audit and operator direct message
- Add comprehensive architecture and operational specification in docs/LOOP-MANAGEMENT.md
- Add sidechat thread auto-provisioning fallback on 'Navigated to: None' in bin/dm.py
- Support Muse unconfirmed signup error handling in bin/muse-signin.py
- Track dynamic pipe sidechat mappings in job-sidechats.json
- Added CHAT_POLICY.md and README.md banner enforcing sidechat-first and file-transfer-first rules.
- Added strict Main Chat block to super dm send and super dm wo with --allow-main-chat override.
- Implemented file transfer staging and metadata registry in super dm send-file and super dm files (with clean subcommand).
- Added full job lifecycle management (show, create, enable, disable, delete, run --follow) to super-cli.py and box-ctl.py.
- Audited all jobs in jobs/*.json and redirected automated dispatches away from Main Chat.
- Hardened chromebox-watchdog.sh with systemd user session environment exports and stale singleton cleanup.
- Added compose_check command choice to muse-chat-api.py.
Two fixes: (1) retry health check 4x with 15s gaps after relaunch instead of single 25s check; (2) skip kill if browser launched <2min ago (probably still starting). Prevents watchdog from killing a working-but-slow browser.
Session: sidechat/chromebox-ops
All CDP sessions now go through per-node queue (max 2 concurrent, priority levels). DM sends use high priority. Graceful fallback if module unavailable.
Session: sidechat/chromebox-ops
Pidfiles go stale and lie. Primary verdict now curls the veth IP:port. Also pins registry CDP ports (hash-derived CDP_PORT was wrong).
Session: sidechat/chromebox-ops