Files
box/docs/SIDECHAT-POLICY-TRIAGE.md
T
operator-main 550e30901b fix: placement verification and sidechat policy enforcement
- dm.py: placement-aware verify (verify_placement), checked_uuid logging, placement_mismatch events
- main-chat-watchdog.py: P0 alerts on placement_mismatch
- box-ctl.py: notify routes to sidechat, box policy command
- jobs: ops-audit and pipe-demo use sidechats
- response-harvester.py: chain deduplication

Session: sidechat/ops-restore
2026-10-04 18:29:13 +00:00

216 lines
9.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Sidechat-Only Policy — Failure Triage Runbook
**Policy:** DMs go to side chats by default. Main chat requires explicit opt-in
(`--allow-main-chat` / job JSON `"allow_main_chat": true`). Anything that lands
in main without opt-in is either intentional-by-design (final nudges, escalations)
or a bug.
**Standing access path for everything below:**
`~/workspace/bin/ssh-vm.sh` → VM → `ssh super@100.123.153.75` (bl).
All paths below are on **bl** unless noted.
## 0. Where things live
| What | File | Notes |
|---|---|---|
| DM send log | `/home/super/Projects/NetVM/dm-log.jsonl` | append-only, msg previews truncated to 100 chars |
| Job/sweeper log | `/home/super/Projects/NetVM/job-log.jsonl` | types: `job_failed`, `followup_nudged`, `followup_nudge_failed`, `followup_needs_review`, `followup_escalated` |
| Followup records | `/home/super/Projects/NetVM/followups.json` | bl-side controller |
| DM code | `/home/super/Projects/NetVM/bin/dm.py` | policy gate at line ~287, exit 2 |
| Job dispatch | `/home/super/Projects/NetVM/bin/job-dispatch.py` | main-gate at line ~384, exit 1 |
| Nudge sweeper | `/home/super/Projects/NetVM/bin/followup-sweeper.py` | final-nudge→main at line ~136, escalation→main at ~214 |
| VM controller | `/srv/board/server.py` + `/srv/box/bin/box_requests.py` | `POST /api/box/followups/cancel` (since 2026-10-04) |
| box CLI notify | `/home/super/Projects/NetVM/bin/box-ctl.py` | `act_notify` lines ~610–612 |
Quick health snapshot (run from bl):
```bash
grep -c '"type": "main_chat_blocked"' /home/super/Projects/NetVM/dm-log.jsonl
grep '"type": "followup_nudge_failed"' /home/super/Projects/NetVM/job-log.jsonl | tail -3
python3 /home/super/Projects/NetVM/bin/followup-sweeper.py --dry-run --once 2>&1 | tail -3
```
---
## 1. "My DM didn't send"
**First check — was it blocked by the policy?**
```bash
grep '"main_chat_blocked"' /home/super/Projects/NetVM/dm-log.jsonl | tail -5
# or by your message id:
grep '"id": "<your-msg-id>"' /home/super/Projects/NetVM/dm-log.jsonl
```
**What "blocked" looks like:** a `main_chat_blocked` event with your msg id, and
`dm.py` exits **2** with stderr: *"Use a sidechat target, or pass
--allow-main-chat for explicit main-chat sends."* The gate fires **before** any
browser navigation — nothing touched main.
**What "healthy" looks like:** the log shows `send_start` → `nav_ok` →
`verified` (with a real `thread_uuid`, not null) and `target` is a sidechat name.
**Fix — pick one:**
- Meant main? `dm.py send --agent opm --to <agent> --target main --allow-main-chat "<msg>"`
- Meant a sidechat? `dm.py send --agent opm --to <agent> --target "<sidechat name>" "<msg>"`
**Escalate to the user when:** you passed `--allow-main-chat` and it *still*
blocked, or a sidechat target was refused without a nav attempt (code bug in
the gate). A plain policy block is not an escalation — it's the system working.
---
## 2. "Job failed"
**First check — job-dispatch logs and dry-run:**
```bash
grep '"type": "job_failed"' /home/super/Projects/NetVM/job-log.jsonl | tail -5
python3 /home/super/Projects/NetVM/bin/job-dispatch.py <job-name> --dry-run
```
**What "broken" looks like:** dry-run resolves to `opm/main` (or any
`<agent>/main`) and exits **1** with: *'set "allow_main_chat": true, or pass
--allow-main-chat.'* Known jobs in this state: `ops-audit-step1/2/3`,
`pipe-demo-step1/2` — they default to main because their job JSON has no
sidechat config.
**What "healthy" looks like:** dry-run resolves to a sidechat, e.g.
`opm/heartbeat`, `646/646 tasks`, `pip/646-pip-coord`.
**Fix:** edit `jobs/<name>.json` — add either:
```json
"sidechat": { "create": true, "name_template": "<name>", "reuse_key": "<key>" }
```
or `"dm_target": "<sidechat name>"`. For jobs that genuinely belong in main,
set `"allow_main_chat": true` instead.
**Escalate to the user when:** a job *has* a sidechat config but still resolves
to main — that's a job-dispatch resolution bug, not a config problem.
---
## 3. "Nudge appeared in main"
**By design:** the final nudge (2/2) and the escalation are *intentionally*
routed to main (`followup-sweeper.py` line ~136:
`delivery_target = "main" if is_final else orig_target`; escalation at ~214).
Early nudges (1/2, etc.) go to the followup's original target.
**Check which case you're in:**
```bash
python3 -c "
import json
d = json.load(open('/home/super/Projects/NetVM/followups.json'))
r = d.get('<nudge-ref-dm-id>')
print('target:', r.get('target'), '| nudges:', r.get('nudges_sent'), '/', r.get('nudges_allowed'), '| status:', r.get('status'))
"
```
**What "healthy" looks like:** 1/2 landed in the original sidechat; 2/2 and
escalation in main. `nudges_sent` advanced once per cycle.
**What "broken" looks like:**
- **1/2 in main** → the original target was main, or sidechat provisioning
failed at send time (see §5).
- **Duplicate 2/2 repeats** → the old failure mode where a failed nudge send
froze state. Fixed 2026-10-04: failed sends now increment `failed_sends`,
push the deadline with backoff (5m→80m), log `followup_nudge_failed`, and set
`needs_review: true` after 3 failures.
**Fix for repeats:** check `grep '"type": "followup_nudge_failed"'
job-log.jsonl | tail -5`. If the record has `needs_review: true`, stand it
down: `super followup cancel <dm_id>` on bl (now also cancels the VM-side
record via `POST /api/box/followups/cancel`).
**Escalate to the user when:** repeats keep firing *after* the backoff fix —
that means the sweeper code regressed.
---
## 4. "box notify doesn't work"
**Known issue (decision pending).** `box-ctl.py` `act_notify` (lines ~610–612)
hard-codes `dm.py send … --target main` *without* `--allow-main-chat`, so the
new policy gate exits 2 and the notify fails. It is not a regression elsewhere
— `box notify` was always main-chat.
**Workaround (explicit, until fixed):**
```bash
python3 /home/super/Projects/NetVM/bin/dm.py send --agent opm --to <agent> \
--target <sidechat> --allow-main-chat "<message>"
```
(keep `--allow-main-chat` only if the sidechat target is truly main)
**Decision needed from the user:** route `box notify` to a per-agent sidechat,
or add `--allow-main-chat` inside `act_notify`. Do not silently re-enable main.
---
## 5. "Ghost followup" (never resolves, nudges forever)
**Root cause:** a followup whose `thread_uuid` is `null` with a non-`main`
target can **never** auto-resolve — the harvester only matches on exact
`thread_uuid` or `target == "main"`. The ghost is born at send time: sidechat
autoprovision failed (`sidechat_uuid_capture_failed` in dm-log, nav landed on
the landing page), but the DM still "sent" with placement-blind `verified:true`,
and the followup was registered with `thread_uuid: null`.
**Check:**
```bash
python3 -c "
import json
d = json.load(open('/home/super/Projects/NetVM/followups.json'))
r = d.get('<dm-id>')
print('thread_uuid:', r.get('thread_uuid'), '| target:', r.get('target'), '| status:', r.get('status'), '| job_id:', r.get('job_id'))
"
grep 'sidechat_uuid_capture_failed' /home/super/Projects/NetVM/dm-log.jsonl | grep '<dm-id>'
```
**What "broken" looks like:** `thread_uuid: None`, `target: <something-not-main>`,
`status: pending` (or escalated), nudges exhausted or still firing.
**Partial mitigation (2026-10-04):** new job DMs store `job_id` in the followup
record, and the harvester resolves `[RESULT <job_id>]` replies by job_id — so
ghosts from *new* sends can still resolve on reply. Followups created *before*
the fix have no `job_id` and need manual resolution.
**Fix (manual resolution):**
```bash
cp /home/super/Projects/NetVM/followups.json /home/super/Projects/NetVM/followups.json.bak-$(date +%Y%m%d)
# then set on the ghost record:
# "status": "resolved", "resolved_manually": true,
# "manual_note": "<reason, e.g. work verified done via RESULT reply; ghost from sidechat_uuid_capture_failed>"
```
Back up first — `followups.json` has no undo.
**Escalate to the user when:** *new* ghosts keep appearing after the job_id fix —
that means sidechat provisioning is regressing again, not just old debris.
---
## 6. Baselines — what "normal" looks like
- `main_chat_blocked` events exist in dm-log (count > 0 is **fine** — policy
working as designed). Worry only if a *legitimate* send is blocked (§1).
- `followup_nudge_failed` in job-log: occasional entries are fine (backoff
absorbs them); a *cluster* means a delivery problem — check dm-log for
`failed`/`retry` chains on the same dm id.
- `followup-sweeper.py --dry-run --once` completes with no tracebacks.
- `box fleet-status`: all nodes `proc_alive:true, cdp_ok:true`. (Note: run
`muse-chat-api.py` only via `netvm-exec.sh <node> --` — direct host invocation
hits `127.0.0.1:94xx` which listens only inside the netns, and will
misleadingly report ConnectionRefusedError.)
- `super followup cancel <dm_id>` now stands down **both** controllers
(bl local + VM `dm_followup` record). If a nudge keeps firing after a
cancel, check the other controller's state before anything else.
## 7. When to wake the user
| Signal | Wake? |
|---|---|
| Single policy block (`main_chat_blocked`) | No — working as designed |
| `--allow-main-chat` still blocked | **Yes** — gate bug |
| Job with sidechat config resolves to main | **Yes** — dispatch bug |
| Duplicate nudges after backoff fix | **Yes** — sweeper regression |
| New null-`thread_uuid` ghosts after job_id fix | **Yes** — provisioning regression |
| `box notify` failure | No — known issue, decision already pending |
| One-off failed send / one ghost | No — use the fixes above |