2026-10-04 18:29:13 +00:00
# Sidechat-Only Policy — Failure Triage Runbook
2026-10-05 15:58:37 +00:00
> **Box is the main surface.** All operator work goes through Box (box.muse-dev.online). The web UI, `box` CLI, and agents share the same API endpoints. No UI-only powers.
2026-10-04 18:29:13 +00:00
**Policy: ** DMs go to side chats by default. Main chat requires explicit opt-in
(`--allow-main-chat` / job JSON `"allow_main_chat": true` ). Anything that lands
in main without opt-in is either intentional-by-design (final nudges, escalations)
or a bug.
**Standing access path for everything below: **
`~/workspace/bin/ssh-vm.sh` → VM → `ssh super@100.123.153.75` (bl).
All paths below are on **bl ** unless noted.
## 0. Where things live
| What | File | Notes |
|---|---|---|
| DM send log | `/home/super/Projects/NetVM/dm-log.jsonl` | append-only, msg previews truncated to 100 chars |
| Job/sweeper log | `/home/super/Projects/NetVM/job-log.jsonl` | types: `job_failed` , `followup_nudged` , `followup_nudge_failed` , `followup_needs_review` , `followup_escalated` |
| Followup records | `/home/super/Projects/NetVM/followups.json` | bl-side controller |
| DM code | `/home/super/Projects/NetVM/bin/dm.py` | policy gate at line ~287, exit 2 |
| Job dispatch | `/home/super/Projects/NetVM/bin/job-dispatch.py` | main-gate at line ~384, exit 1 |
| Nudge sweeper | `/home/super/Projects/NetVM/bin/followup-sweeper.py` | final-nudge→main at line ~136, escalation→main at ~214 |
| VM controller | `/srv/board/server.py` + `/srv/box/bin/box_requests.py` | `POST /api/box/followups/cancel` (since 2026-10-04) |
| box CLI notify | `/home/super/Projects/NetVM/bin/box-ctl.py` | `act_notify` lines ~610– 612 |
Quick health snapshot (run from bl):
``` bash
grep -c '"type": "main_chat_blocked"' /home/super/Projects/NetVM/dm-log.jsonl
grep '"type": "followup_nudge_failed"' /home/super/Projects/NetVM/job-log.jsonl | tail -3
python3 /home/super/Projects/NetVM/bin/followup-sweeper.py --dry-run --once 2>& 1 | tail -3
```
---
## 1. "My DM didn't send"
**First check — was it blocked by the policy? **
``` bash
grep '"main_chat_blocked"' /home/super/Projects/NetVM/dm-log.jsonl | tail -5
# or by your message id:
grep '"id": "<your-msg-id>"' /home/super/Projects/NetVM/dm-log.jsonl
```
**What "blocked" looks like: ** a `main_chat_blocked` event with your msg id, and
`dm.py` exits **2 ** with stderr: *"Use a sidechat target, or pass
--allow-main-chat for explicit main-chat sends."* The gate fires **before ** any
browser navigation — nothing touched main.
**What "healthy" looks like: ** the log shows `send_start` → `nav_ok` →
`verified` (with a real `thread_uuid` , not null) and `target` is a sidechat name.
**Fix — pick one: **
- Meant main? `dm.py send --agent opm --to <agent> --target main --allow-main-chat "<msg>"`
- Meant a sidechat? `dm.py send --agent opm --to <agent> --target "<sidechat name>" "<msg>"`
**Escalate to the user when: ** you passed `--allow-main-chat` and it * still *
blocked, or a sidechat target was refused without a nav attempt (code bug in
the gate). A plain policy block is not an escalation — it's the system working.
---
## 2. "Job failed"
**First check — job-dispatch logs and dry-run: **
``` bash
grep '"type": "job_failed"' /home/super/Projects/NetVM/job-log.jsonl | tail -5
python3 /home/super/Projects/NetVM/bin/job-dispatch.py <job-name> --dry-run
```
**What "broken" looks like: ** dry-run resolves to `opm/main` (or any
`<agent>/main` ) and exits **1 ** with: *'set "allow_main_chat": true, or pass
--allow-main-chat.'* Known jobs in this state: `ops-audit-step1/2/3` ,
`pipe-demo-step1/2` — they default to main because their job JSON has no
sidechat config.
**What "healthy" looks like: ** dry-run resolves to a sidechat, e.g.
`opm/heartbeat` , `646/646 tasks` , `pip/646-pip-coord` .
**Fix: ** edit `jobs/<name>.json` — add either:
``` json
"sidechat" : { "create" : true , "name_template" : "<name>" , "reuse_key" : "<key>" }
```
or `"dm_target": "<sidechat name>"` . For jobs that genuinely belong in main,
set `"allow_main_chat": true` instead.
**Escalate to the user when: ** a job * has * a sidechat config but still resolves
to main — that's a job-dispatch resolution bug, not a config problem.
---
## 3. "Nudge appeared in main"
**By design: ** the final nudge (2/2) and the escalation are * intentionally *
routed to main (`followup-sweeper.py` line ~136:
`delivery_target = "main" if is_final else orig_target` ; escalation at ~214).
Early nudges (1/2, etc.) go to the followup's original target.
**Check which case you're in: **
``` bash
python3 -c "
import json
d = json.load(open('/home/super/Projects/NetVM/followups.json'))
r = d.get('<nudge-ref-dm-id>')
print('target:', r.get('target'), '| nudges:', r.get('nudges_sent'), '/', r.get('nudges_allowed'), '| status:', r.get('status'))
"
```
**What "healthy" looks like: ** 1/2 landed in the original sidechat; 2/2 and
escalation in main. `nudges_sent` advanced once per cycle.
**What "broken" looks like: **
- **1/2 in main** → the original target was main, or sidechat provisioning
failed at send time (see §5).
- **Duplicate 2/2 repeats** → the old failure mode where a failed nudge send
froze state. Fixed 2026-10-04: failed sends now increment `failed_sends` ,
push the deadline with backoff (5m→80m), log `followup_nudge_failed` , and set
`needs_review: true` after 3 failures.
**Fix for repeats: ** check `grep '"type": "followup_nudge_failed"'
job-log.jsonl | tail -5` . If the record has `needs_review: true` , stand it
down: `super followup cancel <dm_id>` on bl (now also cancels the VM-side
record via `POST /api/box/followups/cancel` ).
**Escalate to the user when: ** repeats keep firing * after * the backoff fix —
that means the sweeper code regressed.
---
## 4. "box notify doesn't work"
**Known issue (decision pending). ** `box-ctl.py` `act_notify` (lines ~610– 612)
hard-codes `dm.py send … --target main` * without * `--allow-main-chat` , so the
new policy gate exits 2 and the notify fails. It is not a regression elsewhere
— `box notify` was always main-chat.
**Workaround (explicit, until fixed): **
``` bash
python3 /home/super/Projects/NetVM/bin/dm.py send --agent opm --to <agent> \
--target <sidechat> --allow-main-chat "<message>"
```
(keep `--allow-main-chat` only if the sidechat target is truly main)
**Decision needed from the user: ** route `box notify` to a per-agent sidechat,
or add `--allow-main-chat` inside `act_notify` . Do not silently re-enable main.
---
## 5. "Ghost followup" (never resolves, nudges forever)
**Root cause: ** a followup whose `thread_uuid` is `null` with a non-`main`
target can **never ** auto-resolve — the harvester only matches on exact
`thread_uuid` or `target == "main"` . The ghost is born at send time: sidechat
autoprovision failed (`sidechat_uuid_capture_failed` in dm-log, nav landed on
the landing page), but the DM still "sent" with placement-blind `verified:true` ,
and the followup was registered with `thread_uuid: null` .
**Check: **
``` bash
python3 -c "
import json
d = json.load(open('/home/super/Projects/NetVM/followups.json'))
r = d.get('<dm-id>')
print('thread_uuid:', r.get('thread_uuid'), '| target:', r.get('target'), '| status:', r.get('status'), '| job_id:', r.get('job_id'))
"
grep 'sidechat_uuid_capture_failed' /home/super/Projects/NetVM/dm-log.jsonl | grep '<dm-id>'
```
**What "broken" looks like: ** `thread_uuid: None` , `target: <something-not-main>` ,
`status: pending` (or escalated), nudges exhausted or still firing.
**Partial mitigation (2026-10-04): ** new job DMs store `job_id` in the followup
record, and the harvester resolves `[RESULT <job_id>]` replies by job_id — so
ghosts from * new * sends can still resolve on reply. Followups created * before *
the fix have no `job_id` and need manual resolution.
**Fix (manual resolution): **
``` bash
cp /home/super/Projects/NetVM/followups.json /home/super/Projects/NetVM/followups.json.bak-$( date +%Y%m%d)
# then set on the ghost record:
# "status": "resolved", "resolved_manually": true,
# "manual_note": "<reason, e.g. work verified done via RESULT reply; ghost from sidechat_uuid_capture_failed>"
```
Back up first — `followups.json` has no undo.
**Escalate to the user when: ** * new * ghosts keep appearing after the job_id fix —
that means sidechat provisioning is regressing again, not just old debris.
---
## 6. Baselines — what "normal" looks like
- `main_chat_blocked` events exist in dm-log (count > 0 is **fine ** — policy
working as designed). Worry only if a * legitimate * send is blocked (§1).
- `followup_nudge_failed` in job-log: occasional entries are fine (backoff
absorbs them); a * cluster * means a delivery problem — check dm-log for
`failed` /`retry` chains on the same dm id.
- `followup-sweeper.py --dry-run --once` completes with no tracebacks.
- `box fleet-status` : all nodes `proc_alive:true, cdp_ok:true` . (Note: run
`muse-chat-api.py` only via `netvm-exec.sh <node> --` — direct host invocation
hits `127.0.0.1:94xx` which listens only inside the netns, and will
misleadingly report ConnectionRefusedError.)
- `super followup cancel <dm_id>` now stands down **both ** controllers
(bl local + VM `dm_followup` record). If a nudge keeps firing after a
cancel, check the other controller's state before anything else.
## 7. When to wake the user
| Signal | Wake? |
|---|---|
| Single policy block (`main_chat_blocked` ) | No — working as designed |
| `--allow-main-chat` still blocked | **Yes ** — gate bug |
| Job with sidechat config resolves to main | **Yes ** — dispatch bug |
| Duplicate nudges after backoff fix | **Yes ** — sweeper regression |
| New null-`thread_uuid` ghosts after job_id fix | **Yes ** — provisioning regression |
| `box notify` failure | No — known issue, decision already pending |
| One-off failed send / one ghost | No — use the fixes above |