Files
box/docs/DOM-APPROVALS.md
T
2026-10-04 04:27:16 +00:00

218 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# DOM: Approval & Permission Dialogs
Companion to `DOM-APPROVALS-SPEC.md` (behavioral contract: detection →
classification → handling → exit codes). This doc is the **DOM surface
reference**: what the dialogs look like in the tree, which selectors find
them, how `check_approvals` works today, and what's fragile.
Related: `DOM-EDGE-STATES.md` §5.2, `DOM-PAGE-STRUCTURE.md` §3,
`DOM-INDEX.md` §4f.
## 1. Observed specimen (real, from `docs/pip-approval-dialog.png`)
Captured on pip's node during onboarding. This is a **Muse platform dialog
rendered inside the chat DOM** (bottom-anchored card over the message
stream) — not a browser-chrome permission bubble.
```
┌─────────────────────────────────────────────────────────────┐
│ 🌐 Allow pip to share information with 34.139.37.135? │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ pip wants to connect to a remote server to complete │ │ (expandable detail, chevron)
│ │ onboarding. ⌄ │ │
│ └─────────────────────────────────────────────────────────┘ │
│ [ Allow once ] [ Always allow this site ] [ Deny ] │
└─────────────────────────────────────────────────────────────┘
```
| Field | Observed value |
|---|---|
| Title text | `Allow pip to share information with 34.139.37.135?` |
| Title pattern | `Allow <agent> to share information with <IP>?` |
| Detail text | `pip wants to connect to a remote server to complete onboarding.` |
| Detail pattern | `<agent> wants to <action> to <purpose>.` (collapsible) |
| Buttons (left→right) | `Allow once` (primary/blue), `Always allow this site`, `Deny` |
| Icon | globe/network glyph preceding the title |
Trigger class: the agent attempted an outbound network action (connecting to
a remote server). Other trigger classes are **unobserved** — geolocation,
camera/mic, notifications, clipboard, and download prompts have not been
captured on fleet nodes yet (see §8 induction recipe).
## 2. Dialog DOM structure (as known)
No `[role="dialog"]` or `[role="alertdialog"]` elements have been observed in
settled-state probes (`DOM-PAGE-STRUCTURE.md` §3.2). The dialog is found
**by text, not by structure** — this is the central fragility of the current
implementation.
Known structural facts:
- The dialog is in-DOM (React-rendered), so `Runtime.evaluate` sees it; no
shadow-DOM piercing has been needed so far.
- Buttons are plain `<button>` elements matched by innerText
(`allow` / `deny` / `block`, case-insensitive).
- Per `AGENTS.md` (2026-10-03): React selectors behave identically in
headful and headless environments, so selectors captured headless apply
to the user's browser too — and vice versa.
Candidate selectors to verify on next live capture (none confirmed yet):
```javascript
'[role="dialog"]',
'[role="alertdialog"]',
'[data-testid*="dialog"]',
'[data-testid*="approval"]',
'[data-testid*="permission"]',
// button-level (confirmed pattern, unconfirmed testids):
'button' // innerText matches /allow once|always allow|deny/i
```
## 3. How `check_approvals` works today
Location: `bin/muse-chat-api.py`, `check_approvals(ws)` (~line 70).
Returns `[(dialog_text, is_trusted, action_taken)]`.
**Detection** (two passes, JS evaluated via CDP `Runtime.evaluate`):
1. If `document.body.innerText` contains both `Allow` and `to share`:
collect elements whose `innerText` contains both and is < 500 chars
(first 3, text sliced to 200 chars).
2. Fallback: if no dialog found yet and ≥ 2 `<button>`s have text matching
`allow`/`deny`/`block`, take the first button's closest `div` text
(200 chars).
**Classification** (Python side):
- Extract IPv4s with `\b\d{1,3}\.\d{1,3}\.\d{1,3}\.\d{1,3}\b`.
- **No IP → ignored entirely** (added 2026-10-04 to fix false positives
where chat content mentioning "allow" blocked sends).
- IP in `TRUSTED_IPS` → trusted → auto-approve. Else → untrusted.
**Handling**:
- Trusted: click the button whose text contains `allow once`, else the
button whose text is exactly `allow`. Records `clicked:<text>` or
`NOTFOUND`.
- Untrusted: no click. Action = `APPROVAL_NEEDED`.
**Trusted set** (`TRUSTED_IPS`):
```python
TRUSTED_IPS = {
"34.139.37.135", # VM (gateway)
"100.123.153.75", # bl (main compute)
"100.81.31.9", # VM tailnet
}
```
**Enforcement points** (from `DOM-APPROVALS-SPEC.md`):
| Command | Behavior |
|---|---|
| `send` | Gates: untrusted dialog → `APPROVAL_NEEDED` on stderr, exit **2**, message NOT sent |
| `upload` | Gates (same as `send`) |
| `messages` | Observes only, never gates |
| `wait` | Re-checks every 5s; untrusted → exit 2 |
| `approvals` | Reports; never gates |
## 4. Deliberate safety properties (do not regress)
1. **Least-privilege auto-approve.** The click finder prefers `allow once`
over `allow`. The observed dialog also offers `Always allow this site` —
the finder can never select it: `"always allow this site"` contains
neither `allow once` as a targeted match… more precisely, the predicate
is `t.includes('allow once') || t === 'allow'` (lowercased), and
`"always allow this site" === "allow"` is false while
`"always allow this site".includes("allow once")` is false. Persistent
grants are never auto-clicked. Keep it that way.
2. **Fail closed on untrusted IPs.** Unknown IP → exit 2, human decides.
3. **No click without a dialog.** Buttons are only clicked after a dialog
was positively identified.
## 5. Known gaps & fragilities (input to the approvals work)
1. **Text-only detection, no structural anchor.** If the platform rewords
the dialog (`Allow … to share` → `Grant … access`), detection silently
stops working. Next live capture should confirm or deny
`[role="dialog"]` / `data-testid` hooks and prefer them when present.
2. **False-negative trade (2026-10-04).** Dialogs without an IPv4 are now
*ignored*, not escalated. A real non-IP prompt (e.g. `Allow pip to
share your location?`, camera/mic) would pass silently. Consider:
escalate-on-unknown-pattern vs. ignore — currently ignore. This is the
largest open semantics question.
3. **IP regex is loose.** `\d{1,3}` matches `999.999.999.999` and version
strings like `1.2.3.4` in chat content. A stricter octet check
(0–255) plus context (is the IP inside the dialog element, not just
the 200-char slice?) would cut false trust.
4. **Full-DOM scan per call.** `document.querySelectorAll('*')` with an
`innerText` filter on every element runs on each `send`. Fine today;
scope to overlay/portal regions if it ever shows up in profiles.
5. **`.click()` vs trusted input.** `DOM-INDEX.md` notes Radix controls
often need trusted CDP pointer events, not JS `.click()`. If the
dialog buttons are Radix-based, auto-approve may silently no-op
(records `clicked:` but nothing happens). Verify with a trusted-IP
dialog: does the dialog actually dismiss after auto-approve?
6. **Fallback ordering quirk.** The ≥2-buttons fallback only runs when
pass 1 found nothing. A page with both a share-dialog and unrelated
allow/deny buttons elsewhere takes the share-dialog path — correct —
but a non-share permission dialog on a page that *also* contains the
words "Allow"/"to share" somewhere in chat content could misattribute.
The <500-char element filter mitigates; not proven.
7. **Check-then-act race.** `send` checks once, then sends. A dialog
appearing between check and send is unhandled until the next command.
`wait`'s 5s re-check is the only mitigation.
8. **No deny path.** Automation can approve (trusted) or escalate
(untrusted) but never *deny*. A spurious trusted-IP dialog can only be
approved or left for a human. If denial becomes a needed primitive,
it must be an explicit, logged, human-gated action — never automatic.
9. **Truncation.** Dialog text is sliced to 200 chars at capture, 80 in
the action record. Enough for the IP check; lossy for forensics.
Consider logging the full dialog text to the audit trail on
`APPROVAL_NEEDED`.
## 6. Induction recipe (to reproduce an approval state)
> **Not executed by the documenting subagent** (no live-browser
> capability). Run by a browser-capable operator. **Fleet nodes on bl
> only — never the user's personal browser.**
**Option A — browser permission prompt (deterministic):**
1. Pick a node (pip is a good candidate; CDP 9420 in `warp-pip`).
2. CDP `Page.navigate` to `https://permission.site/` (Google's benign
permission-test page; the established navigate pattern is in
`bin/chat-state-check.py:229` / `bin/meta-acct.py:46`).
3. Trigger one permission (Notification is the least invasive).
4. Capture: full dialog DOM via `Runtime.evaluate` — record tag names,
classes, `data-testid`s, `role`s, button texts and order, and whether
the dialog lives in a portal/root container vs. inline.
5. Clean up: click **Deny**. Confirm the dialog dismisses and the chat
page is unaffected. Navigate back to `https://muse.ai/` and verify
the agent session is intact (`messages` returns content).
**Option B — platform approval via chat (uses existing tooling):**
1. Through the chat API, ask the agent to perform an action that needs
a platform grant (e.g. "download this URL", "share your screen").
The specimen in §1 arose exactly this way (outbound connection
during onboarding).
2. Less deterministic than Option A; useful for capturing *platform*
(not browser-chrome) dialog variants.
**Capture checklist for either option:**
- [ ] Screenshot (like `docs/pip-approval-dialog.png`)
- [ ] `outerHTML` of the dialog container (first 2KB)
- [ ] Confirm/deny: `[role="dialog"]`, `[role="alertdialog"]`, `data-testid`
- [ ] Button list in DOM order with exact innerText
- [ ] Whether `check_approvals` detects it (run `muse-chat-api.py approvals`)
- [ ] Auto-approve behavior if trusted IP (does the dialog actually dismiss?)
**Safety rules:** deny/dismiss when done; never click allow on anything
untrusted; never run induction on the user's personal browser; leave the
node's chat session exactly as found.
## 7. Cross-references
- `DOM-APPROVALS-SPEC.md` — the behavioral contract (detection,
classification, handling, exit codes, operator flow). Read first.
- `DOM-PAGE-STRUCTURE.md` §3 — earlier writeup of `check_approvals`
mechanics and selector patterns.
- `DOM-EDGE-STATES.md` §5.2 — approval dialogs as the one real "banner"
class; false-positive traps for text matching.
- `DOM-INDEX.md` §4f — quick-reference checklist entry.
- `bin/muse-chat-api.py` — `check_approvals(ws)` (~line 70),
`TRUSTED_IPS` (line 47), enforcement in `send`/`upload`/`wait`.
- `docs/pip-approval-dialog.png` — the §1 specimen screenshot.