Files
box/docs/DOM-APPROVALS-SPEC.md
T

3.9 KiB

DOM Headless Approvals Spec (muse.ai automation)

Box is the main surface. All operator work goes through Box (box.muse-dev.online). The web UI, box CLI, and agents share the same API endpoints. No UI-only powers.

Overview

When headless automation (muse-chat-api.py via CDP) drives a muse.ai session, the browser may surface permission/confirmation dialogs that block the flow (e.g. "Allow pip to share information with 34.139.37.135?"). This spec defines how the automation detects, classifies, and handles those dialogs. Principle (from INFRA.md): the chat IS the approval interface — no file-based queue; the automation signals when stuck and the operator resolves it in conversation.

Definitions

  • Approval dialog: any in-DOM permission/confirmation prompt that gates the automation's next action.
  • Trusted origin: an IP in TRUSTED_IPS — our own infrastructure, where auto-approval is safe. Current set:
    • 34.139.37.135 — VM (gateway)
    • 100.123.153.75 — bl (main compute)
    • 100.81.31.9 — VM tailnet
  • APPROVAL_NEEDED: the escalation signal. Printed to stderr as APPROVAL_NEEDED: <dialog summary>, process exits with code 2.

Detection

check_approvals(ws) evaluates in the page DOM:

  1. Body text containing both Allow and to share → permission prompt. Narrows to elements whose innerText contains both and is < 500 chars.
  2. Two or more buttons whose text includes allow, deny, or block → likely permission dialog; captures the closest container's text.

Returns a list of (dialog_text, is_trusted, action_taken).

Classification

Extract IPv4 addresses from the dialog text. The dialog is trusted iff any extracted IP is in TRUSTED_IPS. Dialogs with no recognizable IP are untrusted (fail closed).

Handling

  • Trusted: auto-approve by clicking the button whose text contains allow once, else the button whose text is exactly allow. Records clicked:<button text>.
  • Untrusted: do NOT click. Return APPROVAL_NEEDED; the calling command prints the summary to stderr and exits 2.

Enforcement points

Command Behavior
send Checks approvals first. Untrusted dialog → APPROVAL_NEEDED, exit 2, message NOT sent.
messages Non-blocking check (observes, does not gate).
wait Re-checks every 5s during the wait. Untrusted dialog → APPROVAL_NEEDED, exit 2.
approvals Reports pending dialogs with trusted/action status. Never gates.

Exit-code contract

Code Meaning
0 Done.
1 Failed (bad args, no page, send error, …).
2 APPROVAL_NEEDED — human input required; safe to retry after resolution.

Operator flow

  1. Automation exits 2 with APPROVAL_NEEDED: <details>.
  2. Operator relays the details to the user via chat.
  3. User provides the needed input (e.g. confirms, provides OTP).
  4. Operator re-runs the command; the dialog is either gone or now trusted.

Non-goals (explicitly out of scope)

  • DM content provenance: this spec covers browser dialogs encountered by automation. It does NOT authenticate who authored a DM. A message injected via dm.py send carries no signature; a receiving agent cannot verify the claimed sender from the tooling. (Observed 2026-10-03: pip's agent rightly refused to act on an unsigned DM.) Provenance is a separate spec.
  • Credential approval: Meta account write operations use the human-gated flow in docs/META-ACCOUNTS-API.md, not this spec.

Future work

  • Allowlist dialog types (not just IPs) for finer auto-approval.
  • Structured APPROVAL_NEEDED payloads (JSON) for machine-readable relay.
  • DONE 2026-10-03: dm-level provenance via ssh-keygen signing. dm-sign.sh signs with ssh-keygen -Y sign -n dm (file-based); dm.py send --raw transports the signed block verbatim; dm.py verify-sig verifies against dm-signers/<sender>.pub via ssh-keygen -Y verify.