- bin/dm.py: remove 646 tasks -> 1e75a740 from SIDCHAT_ALIASES. That UUID is stale (not a valid thread on 646 account; SPA redirects elsewhere, causing misdelivery). Alias now falls through to job-sidechats.json/autoprovision. Do NOT re-add hardcoded UUIDs.
- job-sidechats.json: live autoprovisioned mappings for 646-pip and 646 tasks (2026-10-04).
- bin/box-chat-cdp.py: CDP robustness -- drain events until command response, retry switcher lookup while SPA settles, fresh reconnect per retry (transient NO_SWITCHER on pip/opm 2026-10-04).
Session: sidechat/uuid-stale-fix
dm_send already knows the thread UUID at send time (alias resolution, nav URL capture, autoprovision capture) but dropped it: the only stdout channel the board parses carried no thread info, so every auto-logged DM row had thread_uuid null.
Now prints '... SENT and VERIFIED thread=<uuid>' when known, unchanged line otherwise (main-chat sends, UUID-capture failures stay honestly null -- never fabricated). Backward compatible: board _BOX_DM_SENT_RE has no end anchor; bl consumers (job-dispatch.py, siphon-bl.py) use substring matching.
Worktree note: unrelated uncommitted changes remain (job_id followup-correlation hunks in this file, followup-sweeper.py, response-harvester.py, untracked helpers) -- not mine, not staged.
Session: sidechat/box-dms-ui
- Add bin/gravity.py loop diagnostics, reconstruction, and progressive remediation
- Wire hard-break alerting to job-log audit and operator direct message
- Add comprehensive architecture and operational specification in docs/LOOP-MANAGEMENT.md
- Add sidechat thread auto-provisioning fallback on 'Navigated to: None' in bin/dm.py
- Support Muse unconfirmed signup error handling in bin/muse-signin.py
- Track dynamic pipe sidechat mappings in job-sidechats.json
All CDP sessions now go through per-node queue (max 2 concurrent, priority levels). DM sends use high priority. Graceful fallback if module unavailable.
Session: sidechat/chromebox-ops
- New bin/dm-sign.sh: produce signed DMs (ssh-keygen -Y sign, namespace
dm) in the [from:X] [id:Y] wire format that dm.py verify-sig checks.
dm.py referenced it in usage but it never existed.
- dm.py: rewrite stale docstring/argparse (claimed verification was
removed / delivery unconfirmed — it does recipient-side read-back with
3 retries); unify all attribution on [from:X]/[id:Y] (was three
formats: [from X], [X], [from:X]); fix dm_thread double-attribution;
--no-verify kept as a documented no-op; raw mode sends verbatim and
reuses the embedded [id:Y] for the audit log; navigation/send/park
now all target the recipient browser (fixes cross-operator side-chat
sends); drop dead --from-sender flag.
- Register dm-signers/operator-main.pub.
- Round-trip verified: sign (operator-main) -> send --raw via opm
loopback -> read-back SENT+VERIFIED -> verify-sig GOOD.
Root cause: agent-health.service runs Type=oneshot with the default
KillMode=control-group. restart_browser() spawned the replacement
chromium with nohup under the service, so systemd SIGKILLed it the
moment the service exited. Every 5-min tick: FAIL -> restart ->
RECOVERED -> SIGKILL at teardown. opm and pip were permanently dark
and dm.py read masked it (empty output, exit 0).
Fix: launch replacements via systemd-run --user --scope (backgrounded)
so the browser lives in a transient scope outside the service cgroup
and survives teardown. setsid does NOT escape either. Note:
systemd-run --scope waits for the scope even with --no-block
(verified 2026-10-03), hence the backgrounding. Same fix in
chromebox-watchdog.sh (timers currently off).
dm.py: dm_read() now uses run_full() and prints a WARNING to stderr
with rc + last error line instead of failing silently on empty reads.
Trailers: Session: sidechat/opm-blind-fix