Files
box/docs/DOM-APPROVALS.md
T
2026-10-04 04:27:16 +00:00

11 KiB
Raw Blame History

DOM: Approval & Permission Dialogs

Companion to DOM-APPROVALS-SPEC.md (behavioral contract: detection → classification → handling → exit codes). This doc is the DOM surface reference: what the dialogs look like in the tree, which selectors find them, how check_approvals works today, and what's fragile.

Related: DOM-EDGE-STATES.md §5.2, DOM-PAGE-STRUCTURE.md §3, DOM-INDEX.md §4f.

1. Observed specimen (real, from docs/pip-approval-dialog.png)

Captured on pip's node during onboarding. This is a Muse platform dialog rendered inside the chat DOM (bottom-anchored card over the message stream) — not a browser-chrome permission bubble.

┌─────────────────────────────────────────────────────────────┐
│ 🌐  Allow pip to share information with 34.139.37.135?      │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ pip wants to connect to a remote server to complete     │ │  (expandable detail, chevron)
│ │ onboarding.                                          ⌄ │ │
│ └─────────────────────────────────────────────────────────┘ │
│ [ Allow once ]  [ Always allow this site ]  [ Deny ]        │
└─────────────────────────────────────────────────────────────┘
Field Observed value
Title text Allow pip to share information with 34.139.37.135?
Title pattern Allow <agent> to share information with <IP>?
Detail text pip wants to connect to a remote server to complete onboarding.
Detail pattern <agent> wants to <action> to <purpose>. (collapsible)
Buttons (left→right) Allow once (primary/blue), Always allow this site, Deny
Icon globe/network glyph preceding the title

Trigger class: the agent attempted an outbound network action (connecting to a remote server). Other trigger classes are unobserved — geolocation, camera/mic, notifications, clipboard, and download prompts have not been captured on fleet nodes yet (see §8 induction recipe).

2. Dialog DOM structure (as known)

No [role="dialog"] or [role="alertdialog"] elements have been observed in settled-state probes (DOM-PAGE-STRUCTURE.md §3.2). The dialog is found by text, not by structure — this is the central fragility of the current implementation.

Known structural facts:

  • The dialog is in-DOM (React-rendered), so Runtime.evaluate sees it; no shadow-DOM piercing has been needed so far.
  • Buttons are plain <button> elements matched by innerText (allow / deny / block, case-insensitive).
  • Per AGENTS.md (2026-10-03): React selectors behave identically in headful and headless environments, so selectors captured headless apply to the user's browser too — and vice versa.

Candidate selectors to verify on next live capture (none confirmed yet):

'[role="dialog"]',
'[role="alertdialog"]',
'[data-testid*="dialog"]',
'[data-testid*="approval"]',
'[data-testid*="permission"]',
// button-level (confirmed pattern, unconfirmed testids):
'button'  // innerText matches /allow once|always allow|deny/i

3. How check_approvals works today

Location: bin/muse-chat-api.py, check_approvals(ws) (~line 70). Returns [(dialog_text, is_trusted, action_taken)].

Detection (two passes, JS evaluated via CDP Runtime.evaluate):

  1. If document.body.innerText contains both Allow and to share: collect elements whose innerText contains both and is < 500 chars (first 3, text sliced to 200 chars).
  2. Fallback: if no dialog found yet and ≥ 2 <button>s have text matching allow/deny/block, take the first button's closest div text (200 chars).

Classification (Python side):

  • Extract IPv4s with \b\d{1,3}\.\d{1,3}\.\d{1,3}\.\d{1,3}\b.
  • No IP → ignored entirely (added 2026-10-04 to fix false positives where chat content mentioning "allow" blocked sends).
  • IP in TRUSTED_IPS → trusted → auto-approve. Else → untrusted.

Handling:

  • Trusted: click the button whose text contains allow once, else the button whose text is exactly allow. Records clicked:<text> or NOTFOUND.
  • Untrusted: no click. Action = APPROVAL_NEEDED.

Trusted set (TRUSTED_IPS):

TRUSTED_IPS = {
    "34.139.37.135",  # VM (gateway)
    "100.123.153.75", # bl (main compute)
    "100.81.31.9",    # VM tailnet
}

Enforcement points (from DOM-APPROVALS-SPEC.md):

Command Behavior
send Gates: untrusted dialog → APPROVAL_NEEDED on stderr, exit 2, message NOT sent
upload Gates (same as send)
messages Observes only, never gates
wait Re-checks every 5s; untrusted → exit 2
approvals Reports; never gates

4. Deliberate safety properties (do not regress)

  1. Least-privilege auto-approve. The click finder prefers allow once over allow. The observed dialog also offers Always allow this site — the finder can never select it: "always allow this site" contains neither allow once as a targeted match… more precisely, the predicate is t.includes('allow once') || t === 'allow' (lowercased), and "always allow this site" === "allow" is false while "always allow this site".includes("allow once") is false. Persistent grants are never auto-clicked. Keep it that way.
  2. Fail closed on untrusted IPs. Unknown IP → exit 2, human decides.
  3. No click without a dialog. Buttons are only clicked after a dialog was positively identified.

5. Known gaps & fragilities (input to the approvals work)

  1. Text-only detection, no structural anchor. If the platform rewords the dialog (Allow … to share → Grant … access), detection silently stops working. Next live capture should confirm or deny [role="dialog"] / data-testid hooks and prefer them when present.
  2. False-negative trade (2026-10-04). Dialogs without an IPv4 are now ignored, not escalated. A real non-IP prompt (e.g. Allow pip to share your location?, camera/mic) would pass silently. Consider: escalate-on-unknown-pattern vs. ignore — currently ignore. This is the largest open semantics question.
  3. IP regex is loose. \d{1,3} matches 999.999.999.999 and version strings like 1.2.3.4 in chat content. A stricter octet check (0–255) plus context (is the IP inside the dialog element, not just the 200-char slice?) would cut false trust.
  4. Full-DOM scan per call. document.querySelectorAll('*') with an innerText filter on every element runs on each send. Fine today; scope to overlay/portal regions if it ever shows up in profiles.
  5. .click() vs trusted input. DOM-INDEX.md notes Radix controls often need trusted CDP pointer events, not JS .click(). If the dialog buttons are Radix-based, auto-approve may silently no-op (records clicked: but nothing happens). Verify with a trusted-IP dialog: does the dialog actually dismiss after auto-approve?
  6. Fallback ordering quirk. The ≥2-buttons fallback only runs when pass 1 found nothing. A page with both a share-dialog and unrelated allow/deny buttons elsewhere takes the share-dialog path — correct — but a non-share permission dialog on a page that also contains the words "Allow"/"to share" somewhere in chat content could misattribute. The <500-char element filter mitigates; not proven.
  7. Check-then-act race. send checks once, then sends. A dialog appearing between check and send is unhandled until the next command. wait's 5s re-check is the only mitigation.
  8. No deny path. Automation can approve (trusted) or escalate (untrusted) but never deny. A spurious trusted-IP dialog can only be approved or left for a human. If denial becomes a needed primitive, it must be an explicit, logged, human-gated action — never automatic.
  9. Truncation. Dialog text is sliced to 200 chars at capture, 80 in the action record. Enough for the IP check; lossy for forensics. Consider logging the full dialog text to the audit trail on APPROVAL_NEEDED.

6. Induction recipe (to reproduce an approval state)

Not executed by the documenting subagent (no live-browser capability). Run by a browser-capable operator. Fleet nodes on bl only — never the user's personal browser.

Option A — browser permission prompt (deterministic):

  1. Pick a node (pip is a good candidate; CDP 9420 in warp-pip).
  2. CDP Page.navigate to https://permission.site/ (Google's benign permission-test page; the established navigate pattern is in bin/chat-state-check.py:229 / bin/meta-acct.py:46).
  3. Trigger one permission (Notification is the least invasive).
  4. Capture: full dialog DOM via Runtime.evaluate — record tag names, classes, data-testids, roles, button texts and order, and whether the dialog lives in a portal/root container vs. inline.
  5. Clean up: click Deny. Confirm the dialog dismisses and the chat page is unaffected. Navigate back to https://muse.ai/ and verify the agent session is intact (messages returns content).

Option B — platform approval via chat (uses existing tooling):

  1. Through the chat API, ask the agent to perform an action that needs a platform grant (e.g. "download this URL", "share your screen"). The specimen in §1 arose exactly this way (outbound connection during onboarding).
  2. Less deterministic than Option A; useful for capturing platform (not browser-chrome) dialog variants.

Capture checklist for either option:

  • Screenshot (like docs/pip-approval-dialog.png)
  • outerHTML of the dialog container (first 2KB)
  • Confirm/deny: [role="dialog"], [role="alertdialog"], data-testid
  • Button list in DOM order with exact innerText
  • Whether check_approvals detects it (run muse-chat-api.py approvals)
  • Auto-approve behavior if trusted IP (does the dialog actually dismiss?)

Safety rules: deny/dismiss when done; never click allow on anything untrusted; never run induction on the user's personal browser; leave the node's chat session exactly as found.

7. Cross-references

  • DOM-APPROVALS-SPEC.md — the behavioral contract (detection, classification, handling, exit codes, operator flow). Read first.
  • DOM-PAGE-STRUCTURE.md §3 — earlier writeup of check_approvals mechanics and selector patterns.
  • DOM-EDGE-STATES.md §5.2 — approval dialogs as the one real "banner" class; false-positive traps for text matching.
  • DOM-INDEX.md §4f — quick-reference checklist entry.
  • bin/muse-chat-api.py — check_approvals(ws) (~line 70), TRUSTED_IPS (line 47), enforcement in send/upload/wait.
  • docs/pip-approval-dialog.png — the §1 specimen screenshot.