# Sidechat-Only Policy — Failure Triage Runbook **Policy:** DMs go to side chats by default. Main chat requires explicit opt-in (`--allow-main-chat` / job JSON `"allow_main_chat": true`). Anything that lands in main without opt-in is either intentional-by-design (final nudges, escalations) or a bug. **Standing access path for everything below:** `~/workspace/bin/ssh-vm.sh` → VM → `ssh super@100.123.153.75` (bl). All paths below are on **bl** unless noted. ## 0. Where things live | What | File | Notes | |---|---|---| | DM send log | `/home/super/Projects/NetVM/dm-log.jsonl` | append-only, msg previews truncated to 100 chars | | Job/sweeper log | `/home/super/Projects/NetVM/job-log.jsonl` | types: `job_failed`, `followup_nudged`, `followup_nudge_failed`, `followup_needs_review`, `followup_escalated` | | Followup records | `/home/super/Projects/NetVM/followups.json` | bl-side controller | | DM code | `/home/super/Projects/NetVM/bin/dm.py` | policy gate at line ~287, exit 2 | | Job dispatch | `/home/super/Projects/NetVM/bin/job-dispatch.py` | main-gate at line ~384, exit 1 | | Nudge sweeper | `/home/super/Projects/NetVM/bin/followup-sweeper.py` | final-nudge→main at line ~136, escalation→main at ~214 | | VM controller | `/srv/board/server.py` + `/srv/box/bin/box_requests.py` | `POST /api/box/followups/cancel` (since 2026-10-04) | | box CLI notify | `/home/super/Projects/NetVM/bin/box-ctl.py` | `act_notify` lines ~610–612 | Quick health snapshot (run from bl): ```bash grep -c '"type": "main_chat_blocked"' /home/super/Projects/NetVM/dm-log.jsonl grep '"type": "followup_nudge_failed"' /home/super/Projects/NetVM/job-log.jsonl | tail -3 python3 /home/super/Projects/NetVM/bin/followup-sweeper.py --dry-run --once 2>&1 | tail -3 ``` --- ## 1. "My DM didn't send" **First check — was it blocked by the policy?** ```bash grep '"main_chat_blocked"' /home/super/Projects/NetVM/dm-log.jsonl | tail -5 # or by your message id: grep '"id": ""' /home/super/Projects/NetVM/dm-log.jsonl ``` **What "blocked" looks like:** a `main_chat_blocked` event with your msg id, and `dm.py` exits **2** with stderr: *"Use a sidechat target, or pass --allow-main-chat for explicit main-chat sends."* The gate fires **before** any browser navigation — nothing touched main. **What "healthy" looks like:** the log shows `send_start` → `nav_ok` → `verified` (with a real `thread_uuid`, not null) and `target` is a sidechat name. **Fix — pick one:** - Meant main? `dm.py send --agent opm --to --target main --allow-main-chat ""` - Meant a sidechat? `dm.py send --agent opm --to --target "" ""` **Escalate to the user when:** you passed `--allow-main-chat` and it *still* blocked, or a sidechat target was refused without a nav attempt (code bug in the gate). A plain policy block is not an escalation — it's the system working. --- ## 2. "Job failed" **First check — job-dispatch logs and dry-run:** ```bash grep '"type": "job_failed"' /home/super/Projects/NetVM/job-log.jsonl | tail -5 python3 /home/super/Projects/NetVM/bin/job-dispatch.py --dry-run ``` **What "broken" looks like:** dry-run resolves to `opm/main` (or any `/main`) and exits **1** with: *'set "allow_main_chat": true, or pass --allow-main-chat.'* Known jobs in this state: `ops-audit-step1/2/3`, `pipe-demo-step1/2` — they default to main because their job JSON has no sidechat config. **What "healthy" looks like:** dry-run resolves to a sidechat, e.g. `opm/heartbeat`, `646/646 tasks`, `pip/646-pip-coord`. **Fix:** edit `jobs/.json` — add either: ```json "sidechat": { "create": true, "name_template": "", "reuse_key": "" } ``` or `"dm_target": ""`. For jobs that genuinely belong in main, set `"allow_main_chat": true` instead. **Escalate to the user when:** a job *has* a sidechat config but still resolves to main — that's a job-dispatch resolution bug, not a config problem. --- ## 3. "Nudge appeared in main" **By design:** the final nudge (2/2) and the escalation are *intentionally* routed to main (`followup-sweeper.py` line ~136: `delivery_target = "main" if is_final else orig_target`; escalation at ~214). Early nudges (1/2, etc.) go to the followup's original target. **Check which case you're in:** ```bash python3 -c " import json d = json.load(open('/home/super/Projects/NetVM/followups.json')) r = d.get('') print('target:', r.get('target'), '| nudges:', r.get('nudges_sent'), '/', r.get('nudges_allowed'), '| status:', r.get('status')) " ``` **What "healthy" looks like:** 1/2 landed in the original sidechat; 2/2 and escalation in main. `nudges_sent` advanced once per cycle. **What "broken" looks like:** - **1/2 in main** → the original target was main, or sidechat provisioning failed at send time (see §5). - **Duplicate 2/2 repeats** → the old failure mode where a failed nudge send froze state. Fixed 2026-10-04: failed sends now increment `failed_sends`, push the deadline with backoff (5m→80m), log `followup_nudge_failed`, and set `needs_review: true` after 3 failures. **Fix for repeats:** check `grep '"type": "followup_nudge_failed"' job-log.jsonl | tail -5`. If the record has `needs_review: true`, stand it down: `super followup cancel ` on bl (now also cancels the VM-side record via `POST /api/box/followups/cancel`). **Escalate to the user when:** repeats keep firing *after* the backoff fix — that means the sweeper code regressed. --- ## 4. "box notify doesn't work" **Known issue (decision pending).** `box-ctl.py` `act_notify` (lines ~610–612) hard-codes `dm.py send … --target main` *without* `--allow-main-chat`, so the new policy gate exits 2 and the notify fails. It is not a regression elsewhere — `box notify` was always main-chat. **Workaround (explicit, until fixed):** ```bash python3 /home/super/Projects/NetVM/bin/dm.py send --agent opm --to \ --target --allow-main-chat "" ``` (keep `--allow-main-chat` only if the sidechat target is truly main) **Decision needed from the user:** route `box notify` to a per-agent sidechat, or add `--allow-main-chat` inside `act_notify`. Do not silently re-enable main. --- ## 5. "Ghost followup" (never resolves, nudges forever) **Root cause:** a followup whose `thread_uuid` is `null` with a non-`main` target can **never** auto-resolve — the harvester only matches on exact `thread_uuid` or `target == "main"`. The ghost is born at send time: sidechat autoprovision failed (`sidechat_uuid_capture_failed` in dm-log, nav landed on the landing page), but the DM still "sent" with placement-blind `verified:true`, and the followup was registered with `thread_uuid: null`. **Check:** ```bash python3 -c " import json d = json.load(open('/home/super/Projects/NetVM/followups.json')) r = d.get('') print('thread_uuid:', r.get('thread_uuid'), '| target:', r.get('target'), '| status:', r.get('status'), '| job_id:', r.get('job_id')) " grep 'sidechat_uuid_capture_failed' /home/super/Projects/NetVM/dm-log.jsonl | grep '' ``` **What "broken" looks like:** `thread_uuid: None`, `target: `, `status: pending` (or escalated), nudges exhausted or still firing. **Partial mitigation (2026-10-04):** new job DMs store `job_id` in the followup record, and the harvester resolves `[RESULT ]` replies by job_id — so ghosts from *new* sends can still resolve on reply. Followups created *before* the fix have no `job_id` and need manual resolution. **Fix (manual resolution):** ```bash cp /home/super/Projects/NetVM/followups.json /home/super/Projects/NetVM/followups.json.bak-$(date +%Y%m%d) # then set on the ghost record: # "status": "resolved", "resolved_manually": true, # "manual_note": "" ``` Back up first — `followups.json` has no undo. **Escalate to the user when:** *new* ghosts keep appearing after the job_id fix — that means sidechat provisioning is regressing again, not just old debris. --- ## 6. Baselines — what "normal" looks like - `main_chat_blocked` events exist in dm-log (count > 0 is **fine** — policy working as designed). Worry only if a *legitimate* send is blocked (§1). - `followup_nudge_failed` in job-log: occasional entries are fine (backoff absorbs them); a *cluster* means a delivery problem — check dm-log for `failed`/`retry` chains on the same dm id. - `followup-sweeper.py --dry-run --once` completes with no tracebacks. - `box fleet-status`: all nodes `proc_alive:true, cdp_ok:true`. (Note: run `muse-chat-api.py` only via `netvm-exec.sh --` — direct host invocation hits `127.0.0.1:94xx` which listens only inside the netns, and will misleadingly report ConnectionRefusedError.) - `super followup cancel ` now stands down **both** controllers (bl local + VM `dm_followup` record). If a nudge keeps firing after a cancel, check the other controller's state before anything else. ## 7. When to wake the user | Signal | Wake? | |---|---| | Single policy block (`main_chat_blocked`) | No — working as designed | | `--allow-main-chat` still blocked | **Yes** — gate bug | | Job with sidechat config resolves to main | **Yes** — dispatch bug | | Duplicate nudges after backoff fix | **Yes** — sweeper regression | | New null-`thread_uuid` ghosts after job_id fix | **Yes** — provisioning regression | | `box notify` failure | No — known issue, decision already pending | | One-off failed send / one ghost | No — use the fixes above |