Files
box/docs/SIDECHAT-POLICY-TRIAGE.md
T

9.6 KiB
Raw Blame History

Sidechat-Only Policy — Failure Triage Runbook

Box is the main surface. All operator work goes through Box (box.muse-dev.online). The web UI, box CLI, and agents share the same API endpoints. No UI-only powers.

Policy: DMs go to side chats by default. Main chat requires explicit opt-in (--allow-main-chat / job JSON "allow_main_chat": true). Anything that lands in main without opt-in is either intentional-by-design (final nudges, escalations) or a bug.

Standing access path for everything below: ~/workspace/bin/ssh-vm.sh → VM → ssh super@100.123.153.75 (bl). All paths below are on bl unless noted.

0. Where things live

What File Notes
DM send log /home/super/Projects/NetVM/dm-log.jsonl append-only, msg previews truncated to 100 chars
Job/sweeper log /home/super/Projects/NetVM/job-log.jsonl types: job_failed, followup_nudged, followup_nudge_failed, followup_needs_review, followup_escalated
Followup records /home/super/Projects/NetVM/followups.json bl-side controller
DM code /home/super/Projects/NetVM/bin/dm.py policy gate at line ~287, exit 2
Job dispatch /home/super/Projects/NetVM/bin/job-dispatch.py main-gate at line ~384, exit 1
Nudge sweeper /home/super/Projects/NetVM/bin/followup-sweeper.py final-nudge→main at line ~136, escalation→main at ~214
VM controller /srv/board/server.py + /srv/box/bin/box_requests.py POST /api/box/followups/cancel (since 2026-10-04)
box CLI notify /home/super/Projects/NetVM/bin/box-ctl.py act_notify lines ~610–612

Quick health snapshot (run from bl):

grep -c '"type": "main_chat_blocked"' /home/super/Projects/NetVM/dm-log.jsonl
grep '"type": "followup_nudge_failed"' /home/super/Projects/NetVM/job-log.jsonl | tail -3
python3 /home/super/Projects/NetVM/bin/followup-sweeper.py --dry-run --once 2>&1 | tail -3

1. "My DM didn't send"

First check — was it blocked by the policy?

grep '"main_chat_blocked"' /home/super/Projects/NetVM/dm-log.jsonl | tail -5
# or by your message id:
grep '"id": "<your-msg-id>"' /home/super/Projects/NetVM/dm-log.jsonl

What "blocked" looks like: a main_chat_blocked event with your msg id, and dm.py exits 2 with stderr: "Use a sidechat target, or pass --allow-main-chat for explicit main-chat sends." The gate fires before any browser navigation — nothing touched main.

What "healthy" looks like: the log shows send_start → nav_ok → verified (with a real thread_uuid, not null) and target is a sidechat name.

Fix — pick one:

  • Meant main? dm.py send --agent opm --to <agent> --target main --allow-main-chat "<msg>"
  • Meant a sidechat? dm.py send --agent opm --to <agent> --target "<sidechat name>" "<msg>"

Escalate to the user when: you passed --allow-main-chat and it still blocked, or a sidechat target was refused without a nav attempt (code bug in the gate). A plain policy block is not an escalation — it's the system working.


2. "Job failed"

First check — job-dispatch logs and dry-run:

grep '"type": "job_failed"' /home/super/Projects/NetVM/job-log.jsonl | tail -5
python3 /home/super/Projects/NetVM/bin/job-dispatch.py <job-name> --dry-run

What "broken" looks like: dry-run resolves to opm/main (or any <agent>/main) and exits 1 with: 'set "allow_main_chat": true, or pass --allow-main-chat.' Known jobs in this state: ops-audit-step1/2/3, pipe-demo-step1/2 — they default to main because their job JSON has no sidechat config.

What "healthy" looks like: dry-run resolves to a sidechat, e.g. opm/heartbeat, 646/646 tasks, pip/646-pip-coord.

Fix: edit jobs/<name>.json — add either:

"sidechat": { "create": true, "name_template": "<name>", "reuse_key": "<key>" }

or "dm_target": "<sidechat name>". For jobs that genuinely belong in main, set "allow_main_chat": true instead.

Escalate to the user when: a job has a sidechat config but still resolves to main — that's a job-dispatch resolution bug, not a config problem.


3. "Nudge appeared in main"

By design: the final nudge (2/2) and the escalation are intentionally routed to main (followup-sweeper.py line ~136: delivery_target = "main" if is_final else orig_target; escalation at ~214). Early nudges (1/2, etc.) go to the followup's original target.

Check which case you're in:

python3 -c "
import json
d = json.load(open('/home/super/Projects/NetVM/followups.json'))
r = d.get('<nudge-ref-dm-id>')
print('target:', r.get('target'), '| nudges:', r.get('nudges_sent'), '/', r.get('nudges_allowed'), '| status:', r.get('status'))
"

What "healthy" looks like: 1/2 landed in the original sidechat; 2/2 and escalation in main. nudges_sent advanced once per cycle.

What "broken" looks like:

  • 1/2 in main → the original target was main, or sidechat provisioning failed at send time (see §5).
  • Duplicate 2/2 repeats → the old failure mode where a failed nudge send froze state. Fixed 2026-10-04: failed sends now increment failed_sends, push the deadline with backoff (5m→80m), log followup_nudge_failed, and set needs_review: true after 3 failures.

Fix for repeats: check grep '"type": "followup_nudge_failed"' job-log.jsonl | tail -5. If the record has needs_review: true, stand it down: super followup cancel <dm_id> on bl (now also cancels the VM-side record via POST /api/box/followups/cancel).

Escalate to the user when: repeats keep firing after the backoff fix — that means the sweeper code regressed.


4. "box notify doesn't work"

Known issue (decision pending). box-ctl.py act_notify (lines ~610–612) hard-codes dm.py send … --target main without --allow-main-chat, so the new policy gate exits 2 and the notify fails. It is not a regression elsewhere — box notify was always main-chat.

Workaround (explicit, until fixed):

python3 /home/super/Projects/NetVM/bin/dm.py send --agent opm --to <agent> \
  --target <sidechat> --allow-main-chat "<message>"

(keep --allow-main-chat only if the sidechat target is truly main)

Decision needed from the user: route box notify to a per-agent sidechat, or add --allow-main-chat inside act_notify. Do not silently re-enable main.


5. "Ghost followup" (never resolves, nudges forever)

Root cause: a followup whose thread_uuid is null with a non-main target can never auto-resolve — the harvester only matches on exact thread_uuid or target == "main". The ghost is born at send time: sidechat autoprovision failed (sidechat_uuid_capture_failed in dm-log, nav landed on the landing page), but the DM still "sent" with placement-blind verified:true, and the followup was registered with thread_uuid: null.

Check:

python3 -c "
import json
d = json.load(open('/home/super/Projects/NetVM/followups.json'))
r = d.get('<dm-id>')
print('thread_uuid:', r.get('thread_uuid'), '| target:', r.get('target'), '| status:', r.get('status'), '| job_id:', r.get('job_id'))
"
grep 'sidechat_uuid_capture_failed' /home/super/Projects/NetVM/dm-log.jsonl | grep '<dm-id>'

What "broken" looks like: thread_uuid: None, target: <something-not-main>, status: pending (or escalated), nudges exhausted or still firing.

Partial mitigation (2026-10-04): new job DMs store job_id in the followup record, and the harvester resolves [RESULT <job_id>] replies by job_id — so ghosts from new sends can still resolve on reply. Followups created before the fix have no job_id and need manual resolution.

Fix (manual resolution):

cp /home/super/Projects/NetVM/followups.json /home/super/Projects/NetVM/followups.json.bak-$(date +%Y%m%d)
# then set on the ghost record:
#   "status": "resolved", "resolved_manually": true,
#   "manual_note": "<reason, e.g. work verified done via RESULT reply; ghost from sidechat_uuid_capture_failed>"

Back up first — followups.json has no undo.

Escalate to the user when: new ghosts keep appearing after the job_id fix — that means sidechat provisioning is regressing again, not just old debris.


6. Baselines — what "normal" looks like

  • main_chat_blocked events exist in dm-log (count > 0 is fine — policy working as designed). Worry only if a legitimate send is blocked (§1).
  • followup_nudge_failed in job-log: occasional entries are fine (backoff absorbs them); a cluster means a delivery problem — check dm-log for failed/retry chains on the same dm id.
  • followup-sweeper.py --dry-run --once completes with no tracebacks.
  • box fleet-status: all nodes proc_alive:true, cdp_ok:true. (Note: run muse-chat-api.py only via netvm-exec.sh <node> -- — direct host invocation hits 127.0.0.1:94xx which listens only inside the netns, and will misleadingly report ConnectionRefusedError.)
  • super followup cancel <dm_id> now stands down both controllers (bl local + VM dm_followup record). If a nudge keeps firing after a cancel, check the other controller's state before anything else.

7. When to wake the user

Signal Wake?
Single policy block (main_chat_blocked) No — working as designed
--allow-main-chat still blocked Yes — gate bug
Job with sidechat config resolves to main Yes — dispatch bug
Duplicate nudges after backoff fix Yes — sweeper regression
New null-thread_uuid ghosts after job_id fix Yes — provisioning regression
box notify failure No — known issue, decision already pending
One-off failed send / one ghost No — use the fixes above