9.6 KiB
Sidechat-Only Policy — Failure Triage Runbook
Box is the main surface. All operator work goes through Box (box.muse-dev.online). The web UI,
boxCLI, and agents share the same API endpoints. No UI-only powers.
Policy: DMs go to side chats by default. Main chat requires explicit opt-in
(--allow-main-chat / job JSON "allow_main_chat": true). Anything that lands
in main without opt-in is either intentional-by-design (final nudges, escalations)
or a bug.
Standing access path for everything below:
~/workspace/bin/ssh-vm.sh → VM → ssh super@100.123.153.75 (bl).
All paths below are on bl unless noted.
0. Where things live
| What | File | Notes |
|---|---|---|
| DM send log | /home/super/Projects/NetVM/dm-log.jsonl |
append-only, msg previews truncated to 100 chars |
| Job/sweeper log | /home/super/Projects/NetVM/job-log.jsonl |
types: job_failed, followup_nudged, followup_nudge_failed, followup_needs_review, followup_escalated |
| Followup records | /home/super/Projects/NetVM/followups.json |
bl-side controller |
| DM code | /home/super/Projects/NetVM/bin/dm.py |
policy gate at line ~287, exit 2 |
| Job dispatch | /home/super/Projects/NetVM/bin/job-dispatch.py |
main-gate at line ~384, exit 1 |
| Nudge sweeper | /home/super/Projects/NetVM/bin/followup-sweeper.py |
final-nudge→main at line ~136, escalation→main at ~214 |
| VM controller | /srv/board/server.py + /srv/box/bin/box_requests.py |
POST /api/box/followups/cancel (since 2026-10-04) |
| box CLI notify | /home/super/Projects/NetVM/bin/box-ctl.py |
act_notify lines ~610–612 |
Quick health snapshot (run from bl):
grep -c '"type": "main_chat_blocked"' /home/super/Projects/NetVM/dm-log.jsonl
grep '"type": "followup_nudge_failed"' /home/super/Projects/NetVM/job-log.jsonl | tail -3
python3 /home/super/Projects/NetVM/bin/followup-sweeper.py --dry-run --once 2>&1 | tail -3
1. "My DM didn't send"
First check — was it blocked by the policy?
grep '"main_chat_blocked"' /home/super/Projects/NetVM/dm-log.jsonl | tail -5
# or by your message id:
grep '"id": "<your-msg-id>"' /home/super/Projects/NetVM/dm-log.jsonl
What "blocked" looks like: a main_chat_blocked event with your msg id, and
dm.py exits 2 with stderr: "Use a sidechat target, or pass
--allow-main-chat for explicit main-chat sends." The gate fires before any
browser navigation — nothing touched main.
What "healthy" looks like: the log shows send_start → nav_ok →
verified (with a real thread_uuid, not null) and target is a sidechat name.
Fix — pick one:
- Meant main?
dm.py send --agent opm --to <agent> --target main --allow-main-chat "<msg>" - Meant a sidechat?
dm.py send --agent opm --to <agent> --target "<sidechat name>" "<msg>"
Escalate to the user when: you passed --allow-main-chat and it still
blocked, or a sidechat target was refused without a nav attempt (code bug in
the gate). A plain policy block is not an escalation — it's the system working.
2. "Job failed"
First check — job-dispatch logs and dry-run:
grep '"type": "job_failed"' /home/super/Projects/NetVM/job-log.jsonl | tail -5
python3 /home/super/Projects/NetVM/bin/job-dispatch.py <job-name> --dry-run
What "broken" looks like: dry-run resolves to opm/main (or any
<agent>/main) and exits 1 with: 'set "allow_main_chat": true, or pass
--allow-main-chat.' Known jobs in this state: ops-audit-step1/2/3,
pipe-demo-step1/2 — they default to main because their job JSON has no
sidechat config.
What "healthy" looks like: dry-run resolves to a sidechat, e.g.
opm/heartbeat, 646/646 tasks, pip/646-pip-coord.
Fix: edit jobs/<name>.json — add either:
"sidechat": { "create": true, "name_template": "<name>", "reuse_key": "<key>" }
or "dm_target": "<sidechat name>". For jobs that genuinely belong in main,
set "allow_main_chat": true instead.
Escalate to the user when: a job has a sidechat config but still resolves to main — that's a job-dispatch resolution bug, not a config problem.
3. "Nudge appeared in main"
By design: the final nudge (2/2) and the escalation are intentionally
routed to main (followup-sweeper.py line ~136:
delivery_target = "main" if is_final else orig_target; escalation at ~214).
Early nudges (1/2, etc.) go to the followup's original target.
Check which case you're in:
python3 -c "
import json
d = json.load(open('/home/super/Projects/NetVM/followups.json'))
r = d.get('<nudge-ref-dm-id>')
print('target:', r.get('target'), '| nudges:', r.get('nudges_sent'), '/', r.get('nudges_allowed'), '| status:', r.get('status'))
"
What "healthy" looks like: 1/2 landed in the original sidechat; 2/2 and
escalation in main. nudges_sent advanced once per cycle.
What "broken" looks like:
- 1/2 in main → the original target was main, or sidechat provisioning failed at send time (see §5).
- Duplicate 2/2 repeats → the old failure mode where a failed nudge send
froze state. Fixed 2026-10-04: failed sends now increment
failed_sends, push the deadline with backoff (5m→80m), logfollowup_nudge_failed, and setneeds_review: trueafter 3 failures.
Fix for repeats: check grep '"type": "followup_nudge_failed"' job-log.jsonl | tail -5. If the record has needs_review: true, stand it
down: super followup cancel <dm_id> on bl (now also cancels the VM-side
record via POST /api/box/followups/cancel).
Escalate to the user when: repeats keep firing after the backoff fix — that means the sweeper code regressed.
4. "box notify doesn't work"
Known issue (decision pending). box-ctl.py act_notify (lines ~610–612)
hard-codes dm.py send … --target main without --allow-main-chat, so the
new policy gate exits 2 and the notify fails. It is not a regression elsewhere
— box notify was always main-chat.
Workaround (explicit, until fixed):
python3 /home/super/Projects/NetVM/bin/dm.py send --agent opm --to <agent> \
--target <sidechat> --allow-main-chat "<message>"
(keep --allow-main-chat only if the sidechat target is truly main)
Decision needed from the user: route box notify to a per-agent sidechat,
or add --allow-main-chat inside act_notify. Do not silently re-enable main.
5. "Ghost followup" (never resolves, nudges forever)
Root cause: a followup whose thread_uuid is null with a non-main
target can never auto-resolve — the harvester only matches on exact
thread_uuid or target == "main". The ghost is born at send time: sidechat
autoprovision failed (sidechat_uuid_capture_failed in dm-log, nav landed on
the landing page), but the DM still "sent" with placement-blind verified:true,
and the followup was registered with thread_uuid: null.
Check:
python3 -c "
import json
d = json.load(open('/home/super/Projects/NetVM/followups.json'))
r = d.get('<dm-id>')
print('thread_uuid:', r.get('thread_uuid'), '| target:', r.get('target'), '| status:', r.get('status'), '| job_id:', r.get('job_id'))
"
grep 'sidechat_uuid_capture_failed' /home/super/Projects/NetVM/dm-log.jsonl | grep '<dm-id>'
What "broken" looks like: thread_uuid: None, target: <something-not-main>,
status: pending (or escalated), nudges exhausted or still firing.
Partial mitigation (2026-10-04): new job DMs store job_id in the followup
record, and the harvester resolves [RESULT <job_id>] replies by job_id — so
ghosts from new sends can still resolve on reply. Followups created before
the fix have no job_id and need manual resolution.
Fix (manual resolution):
cp /home/super/Projects/NetVM/followups.json /home/super/Projects/NetVM/followups.json.bak-$(date +%Y%m%d)
# then set on the ghost record:
# "status": "resolved", "resolved_manually": true,
# "manual_note": "<reason, e.g. work verified done via RESULT reply; ghost from sidechat_uuid_capture_failed>"
Back up first — followups.json has no undo.
Escalate to the user when: new ghosts keep appearing after the job_id fix — that means sidechat provisioning is regressing again, not just old debris.
6. Baselines — what "normal" looks like
main_chat_blockedevents exist in dm-log (count > 0 is fine — policy working as designed). Worry only if a legitimate send is blocked (§1).followup_nudge_failedin job-log: occasional entries are fine (backoff absorbs them); a cluster means a delivery problem — check dm-log forfailed/retrychains on the same dm id.followup-sweeper.py --dry-run --oncecompletes with no tracebacks.box fleet-status: all nodesproc_alive:true, cdp_ok:true. (Note: runmuse-chat-api.pyonly vianetvm-exec.sh <node> --— direct host invocation hits127.0.0.1:94xxwhich listens only inside the netns, and will misleadingly report ConnectionRefusedError.)super followup cancel <dm_id>now stands down both controllers (bl local + VMdm_followuprecord). If a nudge keeps firing after a cancel, check the other controller's state before anything else.
7. When to wake the user
| Signal | Wake? |
|---|---|
Single policy block (main_chat_blocked) |
No — working as designed |
--allow-main-chat still blocked |
Yes — gate bug |
| Job with sidechat config resolves to main | Yes — dispatch bug |
| Duplicate nudges after backoff fix | Yes — sweeper regression |
New null-thread_uuid ghosts after job_id fix |
Yes — provisioning regression |
box notify failure |
No — known issue, decision already pending |
| One-off failed send / one ghost | No — use the fixes above |