Files
box/docs/SIDECHAT-POLICY-TRIAGE.md
T

218 lines
9.6 KiB
Markdown
Raw Normal View History

# Sidechat-Only Policy — Failure Triage Runbook
> **Box is the main surface.** All operator work goes through Box (box.muse-dev.online). The web UI, `box` CLI, and agents share the same API endpoints. No UI-only powers.
**Policy:** DMs go to side chats by default. Main chat requires explicit opt-in
(`--allow-main-chat` / job JSON `"allow_main_chat": true`). Anything that lands
in main without opt-in is either intentional-by-design (final nudges, escalations)
or a bug.
**Standing access path for everything below:**
`~/workspace/bin/ssh-vm.sh` → VM → `ssh super@100.123.153.75` (bl).
All paths below are on **bl** unless noted.
## 0. Where things live
| What | File | Notes |
|---|---|---|
| DM send log | `/home/super/Projects/NetVM/dm-log.jsonl` | append-only, msg previews truncated to 100 chars |
| Job/sweeper log | `/home/super/Projects/NetVM/job-log.jsonl` | types: `job_failed`, `followup_nudged`, `followup_nudge_failed`, `followup_needs_review`, `followup_escalated` |
| Followup records | `/home/super/Projects/NetVM/followups.json` | bl-side controller |
| DM code | `/home/super/Projects/NetVM/bin/dm.py` | policy gate at line ~287, exit 2 |
| Job dispatch | `/home/super/Projects/NetVM/bin/job-dispatch.py` | main-gate at line ~384, exit 1 |
| Nudge sweeper | `/home/super/Projects/NetVM/bin/followup-sweeper.py` | final-nudge→main at line ~136, escalation→main at ~214 |
| VM controller | `/srv/board/server.py` + `/srv/box/bin/box_requests.py` | `POST /api/box/followups/cancel` (since 2026-10-04) |
| box CLI notify | `/home/super/Projects/NetVM/bin/box-ctl.py` | `act_notify` lines ~610–612 |
Quick health snapshot (run from bl):
```bash
grep -c '"type": "main_chat_blocked"' /home/super/Projects/NetVM/dm-log.jsonl
grep '"type": "followup_nudge_failed"' /home/super/Projects/NetVM/job-log.jsonl | tail -3
python3 /home/super/Projects/NetVM/bin/followup-sweeper.py --dry-run --once 2>&1 | tail -3
```
---
## 1. "My DM didn't send"
**First check — was it blocked by the policy?**
```bash
grep '"main_chat_blocked"' /home/super/Projects/NetVM/dm-log.jsonl | tail -5
# or by your message id:
grep '"id": "<your-msg-id>"' /home/super/Projects/NetVM/dm-log.jsonl
```
**What "blocked" looks like:** a `main_chat_blocked` event with your msg id, and
`dm.py` exits **2** with stderr: *"Use a sidechat target, or pass
--allow-main-chat for explicit main-chat sends."* The gate fires **before** any
browser navigation — nothing touched main.
**What "healthy" looks like:** the log shows `send_start` → `nav_ok` →
`verified` (with a real `thread_uuid`, not null) and `target` is a sidechat name.
**Fix — pick one:**
- Meant main? `dm.py send --agent opm --to <agent> --target main --allow-main-chat "<msg>"`
- Meant a sidechat? `dm.py send --agent opm --to <agent> --target "<sidechat name>" "<msg>"`
**Escalate to the user when:** you passed `--allow-main-chat` and it *still*
blocked, or a sidechat target was refused without a nav attempt (code bug in
the gate). A plain policy block is not an escalation — it's the system working.
---
## 2. "Job failed"
**First check — job-dispatch logs and dry-run:**
```bash
grep '"type": "job_failed"' /home/super/Projects/NetVM/job-log.jsonl | tail -5
python3 /home/super/Projects/NetVM/bin/job-dispatch.py <job-name> --dry-run
```
**What "broken" looks like:** dry-run resolves to `opm/main` (or any
`<agent>/main`) and exits **1** with: *'set "allow_main_chat": true, or pass
--allow-main-chat.'* Known jobs in this state: `ops-audit-step1/2/3`,
`pipe-demo-step1/2` — they default to main because their job JSON has no
sidechat config.
**What "healthy" looks like:** dry-run resolves to a sidechat, e.g.
`opm/heartbeat`, `646/646 tasks`, `pip/646-pip-coord`.
**Fix:** edit `jobs/<name>.json` — add either:
```json
"sidechat": { "create": true, "name_template": "<name>", "reuse_key": "<key>" }
```
or `"dm_target": "<sidechat name>"`. For jobs that genuinely belong in main,
set `"allow_main_chat": true` instead.
**Escalate to the user when:** a job *has* a sidechat config but still resolves
to main — that's a job-dispatch resolution bug, not a config problem.
---
## 3. "Nudge appeared in main"
**By design:** the final nudge (2/2) and the escalation are *intentionally*
routed to main (`followup-sweeper.py` line ~136:
`delivery_target = "main" if is_final else orig_target`; escalation at ~214).
Early nudges (1/2, etc.) go to the followup's original target.
**Check which case you're in:**
```bash
python3 -c "
import json
d = json.load(open('/home/super/Projects/NetVM/followups.json'))
r = d.get('<nudge-ref-dm-id>')
print('target:', r.get('target'), '| nudges:', r.get('nudges_sent'), '/', r.get('nudges_allowed'), '| status:', r.get('status'))
"
```
**What "healthy" looks like:** 1/2 landed in the original sidechat; 2/2 and
escalation in main. `nudges_sent` advanced once per cycle.
**What "broken" looks like:**
- **1/2 in main** → the original target was main, or sidechat provisioning
failed at send time (see §5).
- **Duplicate 2/2 repeats** → the old failure mode where a failed nudge send
froze state. Fixed 2026-10-04: failed sends now increment `failed_sends`,
push the deadline with backoff (5m→80m), log `followup_nudge_failed`, and set
`needs_review: true` after 3 failures.
**Fix for repeats:** check `grep '"type": "followup_nudge_failed"'
job-log.jsonl | tail -5`. If the record has `needs_review: true`, stand it
down: `super followup cancel <dm_id>` on bl (now also cancels the VM-side
record via `POST /api/box/followups/cancel`).
**Escalate to the user when:** repeats keep firing *after* the backoff fix —
that means the sweeper code regressed.
---
## 4. "box notify doesn't work"
**Known issue (decision pending).** `box-ctl.py` `act_notify` (lines ~610–612)
hard-codes `dm.py send … --target main` *without* `--allow-main-chat`, so the
new policy gate exits 2 and the notify fails. It is not a regression elsewhere
— `box notify` was always main-chat.
**Workaround (explicit, until fixed):**
```bash
python3 /home/super/Projects/NetVM/bin/dm.py send --agent opm --to <agent> \
--target <sidechat> --allow-main-chat "<message>"
```
(keep `--allow-main-chat` only if the sidechat target is truly main)
**Decision needed from the user:** route `box notify` to a per-agent sidechat,
or add `--allow-main-chat` inside `act_notify`. Do not silently re-enable main.
---
## 5. "Ghost followup" (never resolves, nudges forever)
**Root cause:** a followup whose `thread_uuid` is `null` with a non-`main`
target can **never** auto-resolve — the harvester only matches on exact
`thread_uuid` or `target == "main"`. The ghost is born at send time: sidechat
autoprovision failed (`sidechat_uuid_capture_failed` in dm-log, nav landed on
the landing page), but the DM still "sent" with placement-blind `verified:true`,
and the followup was registered with `thread_uuid: null`.
**Check:**
```bash
python3 -c "
import json
d = json.load(open('/home/super/Projects/NetVM/followups.json'))
r = d.get('<dm-id>')
print('thread_uuid:', r.get('thread_uuid'), '| target:', r.get('target'), '| status:', r.get('status'), '| job_id:', r.get('job_id'))
"
grep 'sidechat_uuid_capture_failed' /home/super/Projects/NetVM/dm-log.jsonl | grep '<dm-id>'
```
**What "broken" looks like:** `thread_uuid: None`, `target: <something-not-main>`,
`status: pending` (or escalated), nudges exhausted or still firing.
**Partial mitigation (2026-10-04):** new job DMs store `job_id` in the followup
record, and the harvester resolves `[RESULT <job_id>]` replies by job_id — so
ghosts from *new* sends can still resolve on reply. Followups created *before*
the fix have no `job_id` and need manual resolution.
**Fix (manual resolution):**
```bash
cp /home/super/Projects/NetVM/followups.json /home/super/Projects/NetVM/followups.json.bak-$(date +%Y%m%d)
# then set on the ghost record:
# "status": "resolved", "resolved_manually": true,
# "manual_note": "<reason, e.g. work verified done via RESULT reply; ghost from sidechat_uuid_capture_failed>"
```
Back up first — `followups.json` has no undo.
**Escalate to the user when:** *new* ghosts keep appearing after the job_id fix —
that means sidechat provisioning is regressing again, not just old debris.
---
## 6. Baselines — what "normal" looks like
- `main_chat_blocked` events exist in dm-log (count > 0 is **fine** — policy
working as designed). Worry only if a *legitimate* send is blocked (§1).
- `followup_nudge_failed` in job-log: occasional entries are fine (backoff
absorbs them); a *cluster* means a delivery problem — check dm-log for
`failed`/`retry` chains on the same dm id.
- `followup-sweeper.py --dry-run --once` completes with no tracebacks.
- `box fleet-status`: all nodes `proc_alive:true, cdp_ok:true`. (Note: run
`muse-chat-api.py` only via `netvm-exec.sh <node> --` — direct host invocation
hits `127.0.0.1:94xx` which listens only inside the netns, and will
misleadingly report ConnectionRefusedError.)
- `super followup cancel <dm_id>` now stands down **both** controllers
(bl local + VM `dm_followup` record). If a nudge keeps firing after a
cancel, check the other controller's state before anything else.
## 7. When to wake the user
| Signal | Wake? |
|---|---|
| Single policy block (`main_chat_blocked`) | No — working as designed |
| `--allow-main-chat` still blocked | **Yes** — gate bug |
| Job with sidechat config resolves to main | **Yes** — dispatch bug |
| Duplicate nudges after backoff fix | **Yes** — sweeper regression |
| New null-`thread_uuid` ghosts after job_id fix | **Yes** — provisioning regression |
| `box notify` failure | No — known issue, decision already pending |
| One-off failed send / one ghost | No — use the fixes above |