- Upgrade exec-constrained.py with subagent.spawn, thread.list, thread.view, pipeline.run ops
- Grant full ops permissions to all fleet agent identities (646, pip, muse, opm)
- Implement bin/box-relay.sh zero-dependency client supporting bearer and SSH signature auth
- Add fast hybrid gateway path to dm.py for sub-2s verified deliveries
- Fix wait=0 handling in super-cli.py subagent deployments
- Add hourly check-in jobs and scheduler for 646, pip, muse
- Document agent tooling and relay APIs in docs/AGENT-TOOLING.md
Re-applies three workstreams lost when 793d3d7 committed over uncommitted
edits, reconciled against the parallel track's committed dm.py changes:
- followup-sweeper.py: backfill thread_uuid after successful nudge sends;
record final_nudge_target=main on final-nudge routing (C1/C2)
- response-harvester.py: resolve followups on main-chat replies when
final_nudge_target=main (C3); harvest ALL [RESULT] markers per message
- dm.py: pre-send placement gate (fail closed when post-nav URL lacks the
target thread UUID; skips main) — purely additive over 793d3d7+f268d3d
- sidechat_manager.py: wait_for_chat_list() settle-poll for list population
race (sidebar button renders before titles load)
- new: bin/tests/test_followup_fixes.py (25 tests), bin/placement-audit.py,
bin/dm-log-taxonomy.py, bin/session-probe.py,
docs/SIDECHAT-RELIABILITY.md, docs/UUID-ROTATION.md
Verified: 25/25 tests pass, py_compile clean, sweeper/harvester dry-runs clean.
Known limitation: gate catches wrong-placement, not wrong-mapping (false
autoprovision adopting the parked thread needs a creation check).
Thread UUIDs rotate -- heartbeat-opm died twice in one day
(5bd5b350 -> 0077e918 -> dead), dropping self_main_loop digests.
SIDCHAT_ALIASES is now intentionally empty; targets fall through to
job-sidechats.json dynamic mappings then name-based sidechat use with
autoprovision, which self-heals. Also removed stale heartbeat-opm and
pipe-9735f2 entries (dead UUID 0077e918) from job-sidechats.json.
Test: dm.py send to heartbeat resolved via name and SENT+VERIFIED
(thread 5f18476d-8994-49e7-a9e0-4838732363fe).
Session: sidechat/chromebox-ops
- bin/dm.py: remove 646 tasks -> 1e75a740 from SIDCHAT_ALIASES. That UUID is stale (not a valid thread on 646 account; SPA redirects elsewhere, causing misdelivery). Alias now falls through to job-sidechats.json/autoprovision. Do NOT re-add hardcoded UUIDs.
- job-sidechats.json: live autoprovisioned mappings for 646-pip and 646 tasks (2026-10-04).
- bin/box-chat-cdp.py: CDP robustness -- drain events until command response, retry switcher lookup while SPA settles, fresh reconnect per retry (transient NO_SWITCHER on pip/opm 2026-10-04).
Session: sidechat/uuid-stale-fix
dm_send already knows the thread UUID at send time (alias resolution, nav URL capture, autoprovision capture) but dropped it: the only stdout channel the board parses carried no thread info, so every auto-logged DM row had thread_uuid null.
Now prints '... SENT and VERIFIED thread=<uuid>' when known, unchanged line otherwise (main-chat sends, UUID-capture failures stay honestly null -- never fabricated). Backward compatible: board _BOX_DM_SENT_RE has no end anchor; bl consumers (job-dispatch.py, siphon-bl.py) use substring matching.
Worktree note: unrelated uncommitted changes remain (job_id followup-correlation hunks in this file, followup-sweeper.py, response-harvester.py, untracked helpers) -- not mine, not staged.
Session: sidechat/box-dms-ui
- Add bin/gravity.py loop diagnostics, reconstruction, and progressive remediation
- Wire hard-break alerting to job-log audit and operator direct message
- Add comprehensive architecture and operational specification in docs/LOOP-MANAGEMENT.md
- Add sidechat thread auto-provisioning fallback on 'Navigated to: None' in bin/dm.py
- Support Muse unconfirmed signup error handling in bin/muse-signin.py
- Track dynamic pipe sidechat mappings in job-sidechats.json
All CDP sessions now go through per-node queue (max 2 concurrent, priority levels). DM sends use high priority. Graceful fallback if module unavailable.
Session: sidechat/chromebox-ops
- New bin/dm-sign.sh: produce signed DMs (ssh-keygen -Y sign, namespace
dm) in the [from:X] [id:Y] wire format that dm.py verify-sig checks.
dm.py referenced it in usage but it never existed.
- dm.py: rewrite stale docstring/argparse (claimed verification was
removed / delivery unconfirmed — it does recipient-side read-back with
3 retries); unify all attribution on [from:X]/[id:Y] (was three
formats: [from X], [X], [from:X]); fix dm_thread double-attribution;
--no-verify kept as a documented no-op; raw mode sends verbatim and
reuses the embedded [id:Y] for the audit log; navigation/send/park
now all target the recipient browser (fixes cross-operator side-chat
sends); drop dead --from-sender flag.
- Register dm-signers/operator-main.pub.
- Round-trip verified: sign (operator-main) -> send --raw via opm
loopback -> read-back SENT+VERIFIED -> verify-sig GOOD.
Root cause: agent-health.service runs Type=oneshot with the default
KillMode=control-group. restart_browser() spawned the replacement
chromium with nohup under the service, so systemd SIGKILLed it the
moment the service exited. Every 5-min tick: FAIL -> restart ->
RECOVERED -> SIGKILL at teardown. opm and pip were permanently dark
and dm.py read masked it (empty output, exit 0).
Fix: launch replacements via systemd-run --user --scope (backgrounded)
so the browser lives in a transient scope outside the service cgroup
and survives teardown. setsid does NOT escape either. Note:
systemd-run --scope waits for the scope even with --no-block
(verified 2026-10-03), hence the backgrounding. Same fix in
chromebox-watchdog.sh (timers currently off).
dm.py: dm_read() now uses run_full() and prints a WARNING to stderr
with rc + last error line instead of failing silently on empty reads.
Trailers: Session: sidechat/opm-blind-fix