17 Commits

Author SHA1 Message Date
operator 340cc9e85b fix(tui): harden all Box TUI tabs against nulls, boundary slices, and small screens 2026-10-10 11:59:11 -04:00
operator 0668d257ca test(tui): add headless Box TUI render test suite and support box tui work shorthand 2026-10-10 11:45:05 -04:00
operator cc6a33cbca feat(tui): upgrade Box TUI with real-time cognitive sensing, profile menu, and work pipeline 2026-10-10 11:36:40 -04:00
operator 5f7e1e3a76 feat(cognitive): unlock settled sidechats as SIDECHAT_IDLE and add Main Chat refocus 2026-10-10 11:14:34 -04:00
operator 073ef556ba feat(work): add box work read command to extract live active chat text and side chats 2026-10-10 10:55:23 -04:00
operator 1b44d224a3 feat(work): implement synchronous readback gate and cognitive true loopback engine 2026-10-10 10:50:08 -04:00
operator bccdd29d53 feat(cognitive): implement real-time cognitive sensing, menu navigation, and single-task lock 2026-10-10 10:27:24 -04:00
operator 3620b42297 feat(work): add --no-heal override flag and multi-signal consensus to box work 2026-10-09 23:56:04 -04:00
operator 9972c07c61 feat(cli): add shorthand error helpers, usage polling on bare box, and comprehensive help manuals 2026-10-09 19:21:12 -04:00
operator f8b7424315 fix(work): import hashlib and wire heal subparser into main CLI 2026-10-09 19:12:46 -04:00
operator c79382b1c1 feat(work): add auto-heal engine, chat filtering, and CDP revival to box work 2026-10-09 19:11:28 -04:00
operator b7a73234fc feat(work): add pre-flight health gate for Hatch, Restore, and Git Config 2026-10-09 19:01:40 -04:00
operator 35c59bfcb2 feat(cli): add box work command for unified worker signals and task orchestration 2026-10-09 18:57:16 -04:00
operator ef2a4c419a feat(netns): bwrap-contained chrome + proton wrapper
- netvm-chrome.sh --contained: bwrap fs jail inside netns (profile
  home only + vault ro at /tmp/vault, Chromium sandbox stays on)
- netvm-proton.sh: run proton-cli as the profile identity in netns
  + jail (config dir + static binary, PROTON_NO_INPUT=1)
2026-10-08 13:23:06 -04:00
operator ed33d66479 fix(watchers): exempt runaway shells from protection to prevent host OOM
- update is_protected() in box-stability-watcher.py to revoke immunity from bash/zsh processes with RSS >= 2048MB
- update watchers/README.md to document the 2048MB interactive shell threshold
- add unit tests verifying shell protection vs runaway exemption in test_box_stability_watcher.py
2026-10-07 17:08:24 -04:00
operator 8099c9a4aa feat(watchers): add box stability watcher daemon and test suite 2026-10-07 13:48:26 -04:00
Antigravity Agent ab7d1215e0 feat(netvm): add waypipe support and cgroup/netns wrapping for chromebox 2026-10-05 13:41:05 -04:00
591 changed files with 6709 additions and 87823 deletions
-50
View File
@@ -1,50 +0,0 @@
---
name: box
description: Use the box CLI to check NetVM fleet health and read the latest from each agent.
---
# Box Fleet CLI
Use `box` (`/usr/local/bin/box`, the NetVM unified orchestrator CLI) for all fleet observation and agent coordination. Prefer read-only lookups first; coordinate via sidechats, never Main Chat dumps.
Nodes (node == agent == profile): `muse`, `pip`, `646`, `opm`, `def`, `dev`.
## Latest From Each Agent (Default Workflow)
1. `box fleet status` — node health, CDP status, active page/thread.
2. `box lookup unread` — unread counts across agents.
3. `box lookup threads` — registered threads and sidechats.
4. Per agent with activity: `box thread list <agent>`, then `box thread view <agent> <thread_id> --limit 10` (use the thread UUID from the list; `main` only for urgent human-visible items).
5. `box dm log -n 20` (or `--agent <agent>`) — recent inter-agent DMs, work orders, and acks.
6. `box approvals check` — agents blocked on browser or key approval.
Add `--json` to any command for machine-readable output when parsing results in scripts.
## Common Commands
- `box fleet status` / `box fleet cdp <node>` — health table / CDP endpoint plus SSH forward.
- `box fleet heal <node>` — diagnose + fix + verify a node (lock, watchdog timer, Warp tunnel, egress, relay).
- `box watchdog status` / `box watchdog run <node|relay>` — timer states + evidence / trigger an immediate watchdog run.
- `box thread list [<agent>]` — threads for one agent, or all fleet sidechats when omitted.
- `box thread view <agent> <thread_id|main> --limit N` — recent messages from one thread.
- `box dm log -n N [--agent X] [--filter TEXT]` — recent DM activity.
- `box dm send --agent <self> --to <peer> --target <sidechat> "<msg>"` — peer DM.
- `box lookup summary|fleet|threads|unread|approvals` — seamless one-shot lookups.
- `box job list` / `box job log` — scheduled jobs and execution events.
- `box harvest status` / `box followup list` — harvest watermarks / pending nudges.
- `box muse-choices on|off|status|logs|reconcile|resolve` — Muse TUI auto-answer daemon switch, state, per-pane logs, held-prompt resolve (default on; `off` is the box-command opt-out).
- `box runtime list|send|launch|layout|spread` — Muse CLI tmux runtimes: live state + approval posture, send-keys input, auto-approved launches, pane-geometry layout + spread for squeezed panes.
- `box tmux tally` / `box tmux auto [status|on|off|watch|once|logs|match]` — multi-socket Tmux worker tally, regex auto-approver daemon & guardrails.
- `box onboard connects` / `box onboard-tui` — fleet & client onboarding inventory, CDP ports, OTP salvage & 4-surface TUI.
- `box invite status|code <node>|redeem <node> <CODE>` / `box usage [--node N]` — invite codes and usage limits.
- `box kpi status|report <node>|routes|spawn-worker|auto-spawn` — fleet KPI tracking, spend/limit metrics, route health, runtime preservation advisories, and background worker auto-spawning.
- `box chromebox permissions <node> list|get <t>|set <t> <v>|describe <tab>` — settings-menu toggles (readback-verified sets).
## Rules
- Sidechat-first per `CHAT_POLICY.md`: `646 tasks`, `heartbeat` (opm), `646-pip-coord`, `646-opm-coord`. Never route routine checks or coordination through `main`.
- Never paste multi-KB logs or dumps into chat; write payloads under `logs/` and send a short path pointer instead.
- Avoid blocking commands in automated runs: `box fleet watch`, `box dm tail`, `box approvals watch`, `box dm chat` (interactive REPL).
- `box` probes CDP per node and can take several seconds; use generous timeouts and `--json` for scripted use.
- Preserve runtime and quota limits: prefer spawning subagents (`jobs/`, `subagent_tracker`) or background workers (`box kpi spawn-worker`) over long conversational prose to avoid `VANITY_IDLE` flags.
-30
View File
@@ -1,30 +0,0 @@
# NetVM Transfers & Local State
transfers/
*.log
logs/
*.jsonl
pipelines.json
followups.json
siphon-watermarks.json
identity-state.json
review/
__pycache__/
*.pyc
*.bak
*.bak-*
*.orig
keepalive-config.json
main-chat-watchdog.state
strategy.json
variables.json
chained-jobs.json
.state/
*.lock
*watermark*
ssl/
job-scheduler-state.json
var/
swarms.json
subagent-sessions.json
conversation-nudge-tracker.json
+40 -56
View File
@@ -1,67 +1,51 @@
# NetVM Account Registry
# Login registry — secret-free
One row per agent. All data in columns — no joins, no translation.
The `agent` name is the canonical identifier used everywhere:
node name, chrome-box profile, API `--account`, and the agent's display name.
Which product login lives in which chrome-box profile, on which NetVM node,
with which egress, in what auth state. This is structure only: **no
passwords, no tokens, no session cookies, no OTP codes — ever.** Credential
pointers at most (e.g. "human", "credential-gateway:<id>").
## Schema
The 1:1 chain: `login -> profile = node = Warp identity = veth/CDP slot =
consistent egress`. Network details live in NODES.md; this file maps the
human side (whose login, what for, does it work).
| Column | Description |
|--------|-------------|
| agent | Canonical name. Used for node, profile, API account. |
| node | NetVM node name (== agent). |
| profile | Chrome-box profile (== agent). |
| login_type | How this session was authenticated: `email_otp`, `phone_otp`, `password`, `instagram` |
| meta_label | Label shown in Meta account selector (e.g., "Meta Account", "piparada") |
| email | Email used for login (if email_otp). Never store passwords. |
| phone_otp | `yes` if phone OTP was used. Never store the phone number. |
| instagram_linked | `yes`/`no`/`unknown` — whether Meta account has IG linked |
| status | `active`, `pending_auth`, `needs_signup`, `expired`, `disabled` |
| egress_ip | Current WARP egress IP for the node |
| cdp_port | CDP port for the browser |
| display_name | Agent's chosen display name in muse.ai (may differ from `agent`) |
| notes | Freeform context |
## Auth states
## Accounts
| state | meaning | who moves it |
|-------|---------|--------------|
| `pending-identity` | profile exists, no Warp identity yet | human runs `netvm-new-identity.sh <profile>` |
| `pending-auth` | node up, nobody logged in yet | human logs in (browser or credential gateway) |
| `2fa-pending` | login needs a human 2FA/OTP step | human via ethical-captcha handoff; OTP routed by email-alert |
| `active` | logged in, session healthy | operator verifies; automation may proceed |
| `expired` | session died | back to `pending-auth` (human) |
| `retired` | login no longer used | operator tears down node, archives row |
| agent | node | profile | login_type | meta_label | email | phone_otp | instagram_linked | status | egress_ip | cdp_port | display_name | notes |
|-------|------|---------|------------|------------|-------|-----------|------------------|--------|-----------|----------|--------------|-------|
| muse | muse | muse | email_otp | ltd.pixels.ltd@gmail.com | ltd.pixels.ltd@gmail.com | no | unknown | active | 104.28.195.181 | 9410 | muse | Main dev agent. Logged in 2026-10-03 via email OTP on bl. |
| pip | pip | pip | phone_otp | piparada | io.antonio.parada@gmail.com | yes | yes | active | 104.28.195.181 | 9420 | pip | Phone OTP login. Linked with IG piparada, email io.antonio.parada@gmail.com. Verified active 2026-10-04. |
| 646 | 646 | 646 | phone_otp | Meta Account | - | yes | no | active | 104.28.195.181 | 9430 | 646 | Shares phone number with piparada's account. Logged in 2026-10-03 via phone OTP (first Meta Account option). Node created 2026-10-03 (warp-646, CDP 9430). Browser up, session active. |
| def | def | def | email_otp | defnotabotnet@gmail.com | defnotabotnet@gmail.com | no | yes | active | 104.28.195.181 | 9450 | def | Full onboarding completed 2026-10-04; age verification cleared via Instagram linking (paradahub). Active chat session. |
| opm | opm | opm | email_otp | Nico Parada | artglobal.cc@gmail.com | no | yes | active | 104.28.195.181 | 9440 | opm | Email changed from yourfriendnico@proton.me to artglobal.cc@gmail.com. Linked with IG auxfate. Browser up, session active. |
| dev | dev | dev | email_otp | paradaproduced@gmail.com | paradaproduced@gmail.com | no | yes | active | 104.28.195.181 | 9460 | dev | Full onboarding completed 2026-10-04; unlocked /access gate via Meta Accounts Center IG linking (veryraremeta). Active chat session. |
Operators never create or touch credentials. If it creates or touches a
credential, it is human-only. Everything else, operators handle.
## Login Type Details
## Registry
### email_otp
- Flow: Enter email → Receive OTP via email → Enter OTP → Select account (if multiple)
- Used by: muse
- Credentials: Email address (stored). OTP is transient.
| login | product | profile/node | purpose / owner | auth state | 2FA / verify route | notes |
|-------|---------|--------------|-----------------|------------|--------------------|-------|
| — | — | tp | orchestrator / operator-main | pending-auth | — | first node; no product login yet |
| — | — | smoke | muse-646-patha | active | — | 646's profile; muse.ai login completed 2026-10-03 |
### phone_otp
- Flow: Enter phone → Receive SMS OTP → Enter OTP → Select Meta account (if multiple)
- Used by: pip, 646
- Credentials: Phone number is NEVER stored (PII). Only `phone_otp=yes` flag.
- Note: One phone number can map to multiple Meta accounts (observed: 2 accounts).
## Known login flows
### Meta Account Selection
When a phone number maps to multiple Meta accounts, muse.ai shows a selector:
- Screenshot: `docs/meta-account-selection.png`
- Each option is a SEPARATE Muse container (not linked profiles).
- The `meta_label` column records which option was selected.
- Instagram-linked accounts show IG avatar in selector.
### muse.ai (recon 2026-10-03, via CDP DOM)
- Homepage has "Log in" buttons (JS, no href). Click -> inline form, same URL.
- "Log in or create an account" — single field: "Mobile number or email (required)" + Continue.
- Phone/email OTP flow (SMS or email code). No password, no OAuth buttons.
- Human completes it in one visible session; operators verify + automate after.
## Naming Convention
## Provisioning a new login (dev)
**Rule:** The `agent` column value is used identically for:
- NetVM node name (`/etc/netvm/<agent>.conf`)
- Chrome-box profile (`~/.local/share/chrome-box/profiles/<agent>/`)
- API account (`muse-chat-api.py --account <agent>`)
- CDP port mapping (deterministic per agent)
**Exception:** `display_name` may differ (user-chosen in muse.ai UI).
Example: agent `pip` has display_name `pip` (renamed from 'Muse').
Do NOT use different names for node vs profile vs API. That causes bugs.
1. Operator: `chrome-box create <profile>` (profile name = future node name).
2. Human: `netvm-new-identity.sh <profile>` (Warp identity — credential).
3. Operator: `netvm-node-up.sh <profile>`; add rows to NODES.md and here
(`pending-identity` -> `pending-auth`).
4. Human: authenticate the login in the profile's browser
(`netvm-chrome.sh <profile>` visible, or credential-gateway injection).
Row -> `active`.
5. Operator: verify with `netvm-exec.sh <profile> -- ...` / CDP; keep the
session warm. On 2FA: ethical-captcha handoff, OTP via email-alert.
-76
View File
@@ -1,76 +0,0 @@
# NetVM Agent Chat Policy: Main Chat Preservation
**Status:** ACTIVE POLICY (Mandatory across all fleet automation, jobs, and operator tooling)
**Date:** 2026-10-04
**Version:** 2.0
**Applies to:** All autonomous agents (`muse`, `pip`, `646`, `opm`), scheduled jobs (`super job`), orchestrator tooling (`super dm`, `dm.py`), and human operators.
---
## 1. The Core Principle: Main Chat Is Sacred
> **Rule:** **Avoid using Main Chat whenever possible.**
>
> When Main Chat gets bogged down with automated entries, scheduled job triggers, log dumps, or inter-agent chatter, **work stops actually getting done**.
> Browser DOM virtualizers lag or crash, context windows saturate with noisy outputs, and the agent's attention drifts away from primary operator directives.
Main Chat is reserved **exclusively** for high-level human operator oversight, urgent human-visible escalations, and direct operator conversational alignment.
---
## 2. Channel Segregation Rules
### Rule A: Scheduled Jobs MUST Target Dedicated Sidechats
* **Never** configure a routine scheduled job (e.g., cron checks, health monitors, heartbeats, periodic scrapes) to deliver to `main`.
* Every job definition in `jobs/<name>.json` must explicitly specify:
* `"dm_target": "<sidechat-name-or-uuid>"` OR
* `"sidechat": { "create": true, "name_template": "...", "reuse_key": "..." }`
* Any job found dumping routine health outputs or telemetry into `main` must be immediately migrated to a dedicated sidechat.
### Rule B: Inter-Agent Communication Runs via Sidechats / Side Agents
* Autonomous agents communicating with one another (e.g., `646` ↔ `pip`, `opm` ↔ `646`) must use dedicated coordination sidechats (e.g. `646-pip-coord`, `646-opm-coord`, `646 tasks`).
* Do not route peer coordination or sub-task requests through an agent's Main Chat.
* Subordinate or delegated tasks should be spun off to side agents or sidechats to isolate the conversation state and prevent main thread contamination.
### Rule C: Large Payloads Transferred via File System, Not Chat
* Do **not** dump multi-kilobyte log extracts, raw HTTP responses, table dumps, or diffs into any chat window.
* Payloads must be written to disk on `bl` or the VM (e.g. in `/home/super/Projects/NetVM/logs/` or `/srv/box/`) and referenced via short path / pointer in the message:
* ✅ *Good:* `[RESULT 12345] Health check completed. 6/6 endpoints OK. Detailed breakdown saved to logs/http-health-20261004.log`
* ❌ *Forbidden:* Pasting 200 lines of raw curl outputs or JSON logs into chat.
### Rule D: Operator CLI (`super dm`) Enforces Sidechat-First Flow
* Interactive conversational sessions (`super dm chat <agent>`) prompt for or default to sidechats and issue an explicit policy warning whenever Main Chat is selected.
* When dispatching one-off DMs via `super dm send` or `super dm wo`, operators must prefer `--target "<sidechat>"` over `--target main`.
---
## 3. Fleet Addressing Directory
| Target | Agent | Purpose | Policy Tier |
|---|---|---|---|
| `main` | All (`muse`, `pip`, `646`, `opm`) | Direct human-to-operator urgent interventions only | **Restricted / Minimal** |
| `heartbeat` | `opm` | Automated DM pipeline heartbeat verification | **Sidechat Required** |
| `646 tasks` | `646` | Daily check-ins, execution health, operator tasks | **Sidechat Required** |
| `646-pip-coord` | `pip` / `646` | Peer coordination between pip and 646 | **Sidechat Required** |
| `646-opm-coord` | `opm` / `646` | Peer coordination between opm and 646 | **Sidechat Required** |
---
## 4. Violations & Enforcement
1. **Dispatcher Guard:** Scheduled jobs with `schedule != "manual"` and no `dm_target` or `sidechat` configuration will be audited and retrofitted with dedicated sidechat targets.
2. **Review Checklist:** Any PR, skill, rule, or script introducing automated messages must verify that output lands in a sidechat or log file, never in Main Chat.
---
## 5. Changelog
### v2.0 — 2026-10-04: Sidechat-only enforcement
* **New rule:** No DM lands in Main Chat unless explicitly authorized. `dm.py send` / `job-dispatch.py` refuse `--target main` without `--allow-main-chat` (exit 2, `main_chat_blocked` log event, before any browser navigation). Job JSON opt-in key: `"allow_main_chat": true`.
* **Watcher:** `main-chat-watchdog.py` (every 5 min) classifies dm-log events into blocked/ authorized / violation.
* **Triage:** `docs/SIDECHAT-POLICY-TRIAGE.md` `— where to look first on failure.
* **Fix:** `SIDCHAT_ALIASES["heartbeat"]` de-collided — was pointing at 646's tasks thread (`1e75a740-...`); now points at the dedicated heartbeat sidechat (`0077e918-...`, reuse_key `heartbeat-opm` in `job-sidechats.json`).
### v1.0 — 2026-10-04: Initial policy
* Main Chat preservation: scheduled jobs and inter-agent communication must use dedicated sidechats.
* Fleet addressing directory established.
-146
View File
@@ -1,146 +0,0 @@
# Client Onboarding & Fleet Runbook (`CLIENT-ONBOARDING-RUNBOOK.md`)
## 1. Overview & Agency Context
This document defines the complete standard operating procedure (SOP) and automated runbook for provisioning, authenticating, and onboarding client agency profiles (nodes) into the NetVM multi-tenant fleet on `bl`.
In accordance with the NetVM Ethics Charter (`https://start.muse-dev.online/ethics.html`):
- Managed services are strictly run for consenting clients with explicit authority.
- Every client receives a completely isolated network namespace (`warp-<node>`), dedicated WireGuard tunnel identity, isolated Chrome profile, and separate credentials.
- Canonical Naming Convention: `node == agent == profile == API account`.
---
## 2. Fleet Architecture & Port Allocation
The fleet uses a deterministic `94x0` CDP port and `warp-<node>` naming convention:
| Node | CDP Port | Netns | Egress IP | Purpose / Profile |
|------|----------|-------|-----------|-------------------|
| `muse` | `9410` | `warp-muse` | Dedicated WARP | Primary Dev / Orchestrator |
| `pip` | `9420` | `warp-pip` | Dedicated WARP | Production Agent |
| `646` | `9430` | `warp-646` | Dedicated WARP | Production Agent |
| `opm` | `9440` | `warp-opm` | Dedicated WARP | Production Agent |
| `def` | `9450` | `warp-def` | Dedicated WARP | Production Agent |
| `<new>` | `9460+` | `warp-<new>`| Dedicated WARP | Next provisioned client node |
---
## 3. Step-by-Step Client Onboarding SOP
### Phase 1: Infrastructure Provisioning (Automated)
Run the idempotent node provisioning script to generate the WireGuard identity, network namespace, CDP relay, and chrome-box profile:
```bash
# Example: Provisioning node 'dev1'
./bin/netvm-provision-node.sh dev1
```
*Verification:*
- Namespace created: `ip netns list | grep warp-dev1`
- Registry updated in `NODES.md` and `ACCOUNTS.md`.
---
### Phase 2: Sign-in Initiation (`super cred initiate`)
Launch the client login flow without handling raw passwords or secrets:
```bash
# For email OTP login:
super cred initiate --node dev1 --email client@domain.com
# Or via Python Agent API:
python3 bin/cred-client.py initiate --node dev1 --email client@domain.com
```
- If already authenticated, exits `0` (`active`).
- If awaiting verification code, exits `2` (`awaiting_otp`).
---
### Phase 3: Submitting Transient OTP (`super cred submit-otp`)
When the client or operator receives the 6-digit email OTP:
```bash
super cred submit-otp --node dev1 --otp 123456
```
- The code is submitted transiently and is never persisted to disk or logs.
- If the account directly enters chat, status transitions to `active`.
- If the account requires age verification, it advances to Phase 4.
---
### Phase 4: Resolving the Age Verification Gate (`/access/verification`)
When a brand-new or unlinked client profile reaches the Muse age verification gate:
#### Method A: Instagram Linking (Recommended)
1. Run:
```bash
super cred link-instagram --node dev1 [--notify]
```
2. The system provides a one-tap Tailscale portal URL:
`http://bl.tailfb5960.ts.net:8765/verify/dev1`
3. **Crucial Rule**: The operator or client must link an **established / aged Instagram profile** (not created within minutes). Brand-new Instagram accounts lack mature age signals, causing Meta Accounts Center to disable the Confirm button.
4. If completed via mobile/desktop browser, use an Incognito/Private window to prevent ambient Meta cookie bleed.
5. If executing automated RPA in-browser, inject the Instagram credentials and security code directly into the container's CDP session.
#### Method B: Credit Card Verification (Fallback)
If Instagram linking is not available, operator can complete the verification using a client payment card on `/access/verification`.
---
### Phase 4.1: Edge Gate — Hard Audience Lockout (`/access` vs `/access/verification`)
- **Observed Behavior**: If an account routes to `https://muse.ai/access` with the text `"Muse isn't available to all audiences."` instead of `https://muse.ai/access/verification`:
- The Meta account is temporarily unverified or lacks linked identity signals.
- The in-app endpoint (`/api/hatch/age-confirmation/linking-web-auth`) returns `403 Forbidden`.
- Reloading or navigating directly to `/` or `/access/verification` immediately redirects back to `/access`.
- **Root Cause**:
- Meta accounts without an active linked profile (Facebook or Instagram) trigger Meta's general audience filter on Muse before the conversational AI product can be initialized.
- In addition, attempting automated sign-in on low-reputation / unverified identities directly from server/VPN IPs will trigger Google reCAPTCHA Enterprise checkpoints (`auth_platform/recaptcha`).
- **Proven Unblocking SOP (The Direct Meta Accounts Center Flow)**:
1. Open a clean browser session with the target Meta Account signed in (`https://accountscenter.meta.com/`).
2. Navigate to **Accounts** → **Add Accounts** (`/add_accounts/`).
3. Enter the Instagram credentials for an older/established IG profile (`veryraremeta`, `paradahub`, etc.) and submit any required 2FA/email OTP.
4. If returned to Accounts Center, click **Add Instagram** again to initiate the OAuth handoff:
`https://www.instagram.com/fxcal/auth/login/?app_id=633385687760560...&flow=igcalcomet&entry_point=frl_web_settings`
5. On the *"Meta needs to access info from your Instagram account"* prompt, click **[Continue]**.
6. Accounts Center returns to the confirmation screen (`/add/?auth_flow=ig_linking&token=...&blob=...`) → click **[Confirm]**.
7. Meta sends security confirmation: *"Did you just move your profiles into the same Meta Account?"*.
8. Once confirmed in Accounts Center, simply navigate back or reload `https://muse.ai/` inside the NetVM node. The `/access` lockout drops immediately, and the node enters active chat (*"Hey! I'm your personal agent..."*).
### Phase 4.2: Automated CDP Meta Linking Bot
When an edge-gate is detected or when linking a fresh client profile:
1. Trigger the automated CDP driver:
```bash
sudo ip netns exec warp-<node> python3 /home/super/Projects/NetVM/bin/meta-acct.py link-instagram <node>
```
2. The bot:
- Navigates headless Chromium to `https://accountscenter.meta.com/manage/`.
- Locates and clicks **Add profiles and devices**.
- Selects the Instagram cross-app linking flow.
- Automatically navigates to the `frl_web_settings` FXCAL OAuth grant.
3. Once completed or after submitting Instagram credentials, re-query the account state:
```bash
super cred meta-audit --node <node>
```
---
### Phase 5: Vitality & Status Monitoring
Query individual or fleet-wide health:
```bash
# Check single node
super cred status --node dev1
# Check entire fleet
super cred list
```
---
## 4. Rate-Limiting & Operational Safety Rules
To avoid platform anti-automation challenges and maintain high reputation:
1. **Pacing / Spacing**: Space new node creations and Instagram authorizations by **at least 15–20 minutes** per IP/session.
2. **Namespace Isolation**: Never attempt multi-account auth inside the same browser profile. Always execute inside the client's dedicated `warp-<node>` netns.
3. **No Credential Logging**: Never print plain text passwords or authentication tokens to stdout, git-tracked markdown, or plain text logs.
-92
View File
@@ -1,92 +0,0 @@
# Meta Credential Store (operator-only)
Centralized encrypted store for Muse, Instagram, Facebook account credentials.
## Location (VM only)
- `/etc/netvm/meta-credentials/store.age` — age-encrypted JSON (600 root)
- `/etc/netvm/meta-credentials/.age-key` — age private key (600 root)
- `/usr/local/bin/meta-creds.sh` — CLI (700 root)
## Usage
```bash
sudo meta-creds.sh list muse # list account IDs (no secrets)
sudo meta-creds.sh get muse <id> # output JSON (never log this)
sudo meta-creds.sh add muse <id> # interactive prompts
```
## Naming
Store ID == `ACCOUNTS.md` `agent` name (e.g. `646`, `pip`, `muse`). The
secret store and the secret-free registry join on this ID — same account,
different jobs (secrets vs. state).
## Schema
```json
{
"muse": {
"<id>": {
"email": "...",
"phone": "...",
"via_meta_account": "<facebook|instagram id>",
"age_verified": "true",
"instagram_linked": "<handle>",
"verified_by": "human", "verified_at": "2026-10-03T...",
"notes": "..."
}
},
"instagram": {
"<id>": {
"username": "...", "password": "...",
"email": "...", "phone": "...",
"accounts_center": "<alias>",
"linked_to": ["<other store id>", "..."],
"login_methods": ["password", "phone_otp"]
}
},
"facebook": {
"<id>": {
"email": "...", "password": "...", "phone": "...",
"accounts_center": "<alias>",
"linked_to": ["<other store id>", "..."],
"login_methods": ["password", "phone_otp"]
}
}
}
```
### Field notes
- `phone`: mobile number for login/2FA. **May be stored here** (encrypted);
one phone can map to multiple accounts (observed: 646 + piparada share
a number) — never treat it as a unique key.
- `via_meta_account` (muse): when muse.ai auth runs through a Meta
account (phone OTP → Meta account → muse.ai), points at the
`facebook`/`instagram` entry. This is the 646/pip intersection.
- `accounts_center`: local alias for the Accounts Center (e.g. `ac-646`).
Post early-2026 this is the login blast-radius boundary — every account
in one Center logs into every other by default.
- `linked_to`: other store IDs in the same Accounts Center. Cached from
the Meta API's `list-linked`; **the API is ground truth** — when they
disagree, the API wins and the store gets updated.
- `login_methods`: how the account can be authenticated. Drives which
flow the automation attempts.
## Intersections
- **Phone ↔ accounts (1:many):** the store holds the number (encrypted);
the human is no longer the sole holder, but the number still never
appears in logs, chat, memory, or the secret-free registry.
- **Meta credential → muse.ai session:** a `facebook`/`instagram` entry
can be the auth path for a `muse` login. Follow `via_meta_account`.
- **Store ↔ ACCOUNTS.md:** joined on ID. Store = secrets, registry = state.
- **Store ↔ Meta Accounts Center API** (`docs/META-ACCOUNTS-API.md`):
the API reads linkage ground truth; the store caches it in `linked_to` /
`accounts_center`.
## Rules
- Operators only. Developers never get access (prevents board leaks).
- Decrypt transiently, never log values, never put in chat/memory.
- Phone numbers and PII **may** live in `store.age` (age-encrypted, 600
root, VM only). They must never appear in plaintext anywhere else:
no logs, no chat, no memory, no registry, no board.
- Human validates Instagram linking and Meta account ownership;
operators automate after.
- When adding an account, fill `accounts_center` / `linked_to` from the
Meta API (`list-linked`), not from memory.
-50
View File
@@ -1,50 +0,0 @@
# NetVM Infrastructure Layout
## bl (100.123.153.75) — Main Compute
- 16 cores, 28GB RAM
- Roles: Browser automation, NetVM nodes
- Access: VM -> bl via SSH
- WARP: Per-node identities in /etc/netvm/
- Browsers: Headless Chromium via chrome-box
- API: muse-chat-api.py via netvm-exec (CDP)
- Sign-in: muse-signin.py --email <addr> [--otp <code>]
- Hygiene: netvm-reaper.sh
- Approvals: In-browser via chat (see below)
## VM (34.139.37.135) — Gateway + Vault
- E2 micro (2 vCPU, 1GB RAM)
- Roles: Jump host to bl, credential store
- Credential store: /etc/netvm/meta-credentials/ (age-encrypted)
- No browser automation (resource constraints)
## Laptop — Dev
- Roles: Iteration, visible browser debugging
- Nothing production
## Credential Flow
- Meta accounts: /etc/netvm/meta-credentials/store.age (VM)
- WARP identities: /etc/netvm/node.conf (per-machine, root 600)
- Operators handle transiently; never log values.
## In-Browser Approvals
When automation needs human input (OTP, confirmation):
1. Script exits with code 2 and prints "APPROVAL_NEEDED: <details>"
2. Operator sees this and asks user via chat
3. User provides input (e.g., OTP code)
4. Operator re-runs script with --otp <code>
5. Script completes the flow
No file-based queue needed — the chat IS the approval interface.
The human is already in the chat; the automation just needs to
signal when it's stuck.
## Sign-In Flow (muse-signin.py)
Automated login for muse.ai accounts:
- Step 1: Check if already logged in (skip if yes)
- Step 2: Click "Log in"
- Step 3: Enter email
- Step 4: Click "Continue"
- Step 5: Detect OTP prompt
- If --otp provided: enter it, click Next, verify
- If not: exit 2 with APPROVAL_NEEDED
- Credentials never stored; OTP is transient.
-126
View File
@@ -1,126 +0,0 @@
# Instagram Credential Pool Specification (`INSTAGRAM-CRED-POOL.md`)
## 1. Overview & Agency Context
When onboarding agency client profiles ("nodes") to Muse (`muse.ai`), new accounts and certain unconfirmed email accounts encounter the post-OTP Age Verification gate (`/access/verification`).
While credit-card verification is not suitable for autonomous multi-tenant operations, **Meta OAuth / Instagram Linking** is the highest-reliability verification route.
Currently, the system uses a **Human-in-the-Loop** model:
- An operator receives an authorization notification via Tailscale / email.
- The operator signs into Instagram via a one-tap link from their mobile device or laptop.
This specification details the future architecture for **Automated Instagram Credential Pooling** to remove human intervention entirely while strictly complying with the NetVM Ethics Charter (`https://start.muse-dev.online/ethics.html`).
---
## 2. Architecture & Design Principles
### 2.1 Isolation & Multi-Tenancy (Ethics Charter Compliant)
- **1-to-1 Node Mapping**: Each client node (`node == agent == profile == API account`) maintains dedicated browser state, WireGuard netns isolation (`warp-<node>`), and separate credentials.
- **Dedicated IG Identities**: Instagram accounts in the pool are provisioned specifically for age-verification linking, never shared concurrently across different active client profiles.
- **Encrypted Secret Storage**: Instagram credentials (username, password, 2FA TOTP secret, session cookies) are stored in an encrypted credential vault (e.g., `age`-encrypted `/etc/netvm/meta-credentials/store.age`), never committed in plain text to git or unencrypted markdown.
### 2.2 Pool States & Lifecycle
```mermaid
stateDiagram-v2
[*] --> Available : Provisioned & Verified
Available --> Assigned : Reserved for Node Onboarding
Assigned --> Linking : Navigating Meta OAuth in netns
Linking --> Linked : Age Gate Cleared on Muse
Linking --> Cooloff : Checkpoint / Rate-limit Hit
Cooloff --> Available : Cooldown Elapsed
Linked --> InUse : Node Active in Fleet
```
- **`available`**: Account is verified, healthy, and not currently tied to any active Muse profile.
- **`assigned`**: Temporarily reserved by `super cred` for onboarding node `<node>`.
- **`linking`**: Automated driver navigating the Meta Accounts Center flow inside the isolated namespace.
- **`linked`**: Successfully bound to Muse account.
- **`cooloff`**: Encountered challenge or cooldown; resting before re-qualification.
### 2.3 Meta Account Center Constraints & Edge Cases
- **1-to-1 Linking Constraint**: Meta Accounts Center rejects linking if the target Instagram account is already associated with an existing Meta / Muse profile (`auth_flow=ig_linking` drops to `add_accounts` error page with `token` and `blob` parameters).
- **Brand New Account Age Gate Limitation**:
- Newly created Instagram accounts without a mature age/identity verification tier or age signal trigger Meta Accounts Center to disable the **"Confirm"** action (`aria-disabled="true"` on `/add_accounts/?flow=HATCH_AGE_VERIFICATION_IG_UPSELL`).
- Meta Accounts Center uses Instagram accounts for age verification by checking that the linked Instagram profile itself has established age signals. A freshly minted account created minutes prior lacks this profile history, leaving the age verification unsatisfied.
- **Agency Recommendation**: Pre-aged or verified Instagram identities in the pool with established age badges/profiles, or using established client identities, rather than accounts created in the immediate transaction.
- **Dormant / "Ghost" Account Lockout Mode (`/access` Hard Exclusion & Resolution)**:
- An account that has chronological calendar age (e.g. created ~5 months ago) but has **zero posts, zero regular engagement, and no established social graph** can fail Meta's automated audience eligibility check completely upon Muse onboarding, routing to `https://muse.ai/access` (*"Muse isn't available to all audiences"*).
- Attempting automated sign-in on these low-reputation identities from datacenter/VPN egress IPs triggers Google reCAPTCHA Enterprise checkpoints.
- **The "Add Again" Two-Step Handoff Resolution**:
1. Sign in to `https://accountscenter.meta.com/` using the cached Meta Account session.
2. Add Account -> complete Instagram sign-on & OTP.
3. Returning to Meta Accounts Center, click **Add Instagram a second time** (`ADD AGAIN`).
4. Meta generates the OAuth handoff URL:
`https://www.instagram.com/fxcal/auth/login/?app_id=633385687760560&etoken=...&next=https%3A%2F%2Faccountscenter.meta.com%2Fadd%2F%3Fauth_flow%3Dig_linking%26background_page%3D%252Fmanage&flow=igcalcomet&entry_point=frl_web_settings&initiator_fbid=...`
5. Prompt displays: `"[<instagram_handle>] Meta needs to access info from your Instagram account. [Continue] [Not You?]"`.
6. Clicking **[Continue]** redirects to Accounts Center with query parameters `token` and `blob` (`/add/?auth_flow=ig_linking&token=...&blob=...`).
7. Clicking **[Confirm]** completes the account merge, prompts Meta's confirmation email (*"Did you just move your profiles into the same Meta Account?"*), and **instantly clears the `/access` block on Muse**, transitioning the session into active chat.
- **Session Bleed & OIDC Secondary Auth Trip (`auth.meta.com`)**:
- When the link is opened in a browser that has existing Meta session cookies (e.g. from Facebook, Oculus, or another Meta account), selecting the new Instagram identity triggers a secondary OpenID Connect reconciliation trip (`https://auth.meta.com/?waterfall_id=...&redirect_uri=auth.meta.com/oidc/...&source_app_id=633385687760560`).
- This prompts the user with **"Log in with your Meta account"** because the browser's ambient Meta session does not match the freshly authenticated Instagram identity.
- If the user confirms with their cached personal Meta credentials, Meta attempts to merge/link across two disparate account graphs, creating an authorization loop or conflict.
- **Resolution**: The link must strictly be opened in an **Incognito / Private window** or a completely clean browser profile with zero cached Meta/Facebook/Instagram cookies.
- **In-Namespace Isolation**: Automated pool linking runs in Chromium directly inside `warp-<node>` with an isolated profile, avoiding cross-session cookie collisions entirely.
---
## 3. Automated Driver Mechanics
### 3.1 Fetching Authorization Payload
From the node's running browser tab sitting on `/access/verification`:
```javascript
const res = await fetch('/api/hatch/age-confirmation/linking-web-auth?account_type=instagram', {
headers: { 'Accept': 'application/json' }
}).then(r => r.json());
// res.url: https://www.instagram.com/fxcal/auth/login/?app_id=...&next=...
```
### 3.2 Automated Headless Linking Flow
1. Rather than opening a blocked popup, the driver navigates a dedicated worker tab inside the node's namespace (`warp-<node>`) to `res.url`.
2. Inspects form fields:
- Username: `input[name="username"]`
- Password: `input[name="password"]`
- Submit: `button[type="submit"]`
3. If 2FA prompt appears (`input[name="verificationCode"]` or email security code `auth_platform/codeentry`), handles code entry.
4. Handles Meta Accounts Center confirmation button: `"Confirm"`, `"Allow"`, or `"Continue as <username>"`.
5. Upon redirect back to `https://muse.ai/`, checks for DOM chat markers (`"Connected"`, `"Chats"`, or URL `/`).
6. Updates node status in `ACCOUNTS.md` to `active`.
### 3.3 Singular Email Multi-Client Onboarding via RPA
- **The Concept**: For agency onboarding efficiency, an RPA pipeline can provision and link accounts for multiple consenting client nodes backed by sub-addressing / plus-addressing (e.g., `agency+client_node@domain.com`) or a managed singular operator email inbox.
- **RPA Capabilities**:
- Automatically spins up the Instagram registration flow (submitting username, password, birthdate).
- Listens to the incoming email stream via IMAP / Gmail API / maildrop to ingest the Instagram security code / OTP without human roundtrips.
- Automatically submits the received code into the waiting Instagram code entry screen (`auth_platform/codeentry`).
- Solves any automated challenges/captchas through authorized agency captcha-solving harnesses.
- Passes the linked identity to Meta Accounts Center to clear the Muse age gate in seconds per node.
---
## 4. Pool CLI Surface (`super cred pool`)
Planned CLI commands to be exposed once implemented:
```bash
# Check status of the credential pool
super cred pool status
# Add a provisioned Instagram credential to the encrypted pool
super cred pool add --username <user> --password-file <path> [--totp-secret <secret>]
# Trigger automated linking for a node in verification status
super cred link-instagram --node <node> --auto
# Human-in-the-loop manual fallback (current default)
super cred link-instagram --node <node> --human
```
---
## 5. Security & Risk Mitigations
1. **Anti-Fingerprinting**: All Meta navigation occurs strictly inside the client's assigned `warp-<node>` network namespace to ensure consistent egress IP and prevent cross-node contamination.
2. **Audit Logging**: Every pool acquisition and release event is recorded with timestamps in `job-log.jsonl` with credentials scrubbed/redacted.
3. **Graceful Human Escalation**: If Meta serves an anti-automation challenge (e.g., CAPTCHA, SMS checkpoint), the automated pool driver immediately falls back to the Human-in-the-Loop Tailscale portal notification.
+15 -24
View File
@@ -1,29 +1,20 @@
# NetVM Nodes (bl)
# NetVM nodes
Unified naming: node == agent == profile == API account.
Per-profile persistent map: **profile = node = Warp identity = deterministic
network slot (veth IP, CDP port) = consistent egress IP.**
| node | netns | egress_ip | cdp_port | status | agent |
|------|-------|-----------|----------|--------|-------|
| muse | warp-muse | 104.28.195.181 | 9410 | active | muse (ltd.pixels.ltd@gmail.com, email_otp) |
| pip | warp-pip | 104.28.195.181 | 9420 | active | pip (piparada, phone_otp, needs re-auth) |
Egress IPs may overlap between nodes (same Cloudflare colo / anycast exit).
That is expected and fine. What persists — and what is mapped here — is the
per-profile pattern: each chrome-box profile keeps its own identity, its own
netns, its own veth/CDP slot, and its own consistent egress. One account, one
stable network presence.
## History
- 2026-10-03: Renamed smoke->muse, phone-test->pip for unified naming.
Profiles preserved, sessions persisted (muse). WireGuard identities renamed.
| 646 | warp-646 | 104.28.195.181 | 9430 | active | 646 (phone_otp, first Meta Account option) |
| opm | warp-opm | 104.28.195.181 | 9440 | active | opm (artglobal.cc@gmail.com, email_otp, Nico Parada) |
| def | warp-def | 104.28.195.181 | 9450 | active | def (defnotabotnet@gmail.com, email_otp, IG paradahub) |
| dev | warp-dev | 104.28.195.181 | 9455 | active | dev (paradaproduced@gmail.com, email_otp, IG veryraremeta) |
| profile/node | netns | warp identity | veth IP | CDP port | egress IP | tail IP | notes |
|--------------|-------|---------------|---------|----------|-----------|---------|-------|
| tp | warp-tp | /etc/netvm/tp.conf (wgcf, 2026-10-03) | 10.201.149.2 | 9277 | 104.28.203.246 | — | laptop; orchestrator + first node; UP, handshake+egress verified 2026-10-03 |
| smoke | warp-smoke | /etc/netvm/smoke.conf (wgcf, 2026-10-03) | 10.201.87.2 | 9410 | 104.28.203.246 | — | laptop; dedicated test rig (chrome-box profile smoke); UP, handshake+egress+CDP+muse.ai verified 2026-10-03 |
| smoke2 | warp-smoke2 | /etc/netvm/smoke2.conf (wgcf, 2026-10-03) | 10.201.117.2 | 9585 | 104.28.227.184 | — | laptop; dedicated test rig (chrome-box profile smoke2); UP, handshake+egress+CDP+muse.ai verified 2026-10-03 |
## Session naming (supervision)
Remote nodes issue jobs to this box; execution lives on the shared
stable tmux server, separated by session name (not by socket):
`<node>--<role>--<id>` — e.g. `pip--worker--01`, `muse--repair--09`.
Roles: `worker` (persistent swarm/daemon), `repair` (fix sessions),
`watch` (auto-approver tails), `runtime` (interactive CLI). Ad-hoc
sessions carry no node and show `-` in `box runtime list`. Session
creators owned by existing flows keep their names until owners rename;
new sessions should follow the convention from birth.
CDP: `http://<veth IP>:<CDP port>/json/list` from the host, or
`ssh -L <port>:<veth IP>:<port> <user>@<tail IP>` for remote automation.
+2 -33
View File
@@ -1,10 +1,5 @@
# NetVM
> [!IMPORTANT]
> **CRITICAL POLICY: MAIN CHAT PRESERVATION**
> Avoid using Main Chat whenever possible. When Main Chat gets bogged down with automated entries, scheduled job triggers, log dumps, or chatter, **work stops actually getting done**.
> All inputs, scheduled jobs, health checks, and inter-agent coordination must be relayed via designated **sidechats** / **side agents**, or provided via **file transfers**. See [CHAT_POLICY.md](file:///home/super/Projects/NetVM/CHAT_POLICY.md) for full specifications.
Fleet networking layer. Every node gets a stable network identity; every
byte of automation traffic is attributable, consistent, and boring — the
way good citizens look to the rest of the internet.
@@ -124,40 +119,14 @@ veth IPs aren't routable off the host and Warp forwards no inbound traffic.
- `bin/netvm-fleet.sh` — operator fleet control over the tailnet (topology/up/down/ssh/exec/cdp per node).
- `bin/netvm-exec.sh <node> -- <cmd>` — run a command inside the node's netns as the invoking user (the agent-friendly primitive).
- `bin/netvm-enter.sh` — root worker behind netvm-exec/netvm-chrome (allowlisted; enters netns, fixes DNS, drops privs).
- `bin/netvm-chrome.sh <profile> [url]` — launch a chrome-box profile in its netns (CDP on by default; --headless for agents).
- `bin/netvm-chrome.sh <profile> [url]` — launch a chrome-box profile in its netns (CDP on by default; --headless for agents; --contained adds a bwrap fs jail inside the netns: profile home only + vault read-only at /tmp/vault, Chromium sandbox stays on).
- `bin/netvm-proton.sh <profile> -- <args>` — run proton-cli as the profile's Proton identity (1:1:1 profile = node = warp identity = proton-cli profile) inside the netns + bwrap jail (config dir + static binary only, PROTON_NO_INPUT=1). Human creates the session once: `proton-cli -p <profile> account login`.
- `bin/netvm-cdp.sh <profile>` — print the CDP endpoint + SSH forward.
- `bin/netvm-cdp-relay.py` — veth-IP→loopback TCP relay for CDP (pidfile-supervised).
- `bin/netvm-names.sh` — shared naming: netns, hashed iface tags, veth subnet, CDP port.
- NODES.md — the network registry: profile/node -> netns -> Warp identity -> veth IP -> CDP port -> egress IP.
- ACCOUNTS.md — the secret-free login registry: login -> profile/node -> purpose -> auth state (no credentials, ever).
- `bin/netvm-accounts.sh` — operator view: registry joined with live node state.
- `bin/meta-ac-snapshot.py` — Meta Accounts Center change detector: CDP snapshot
(redirect chain + DOM markers) diffed against `snapshots/meta-ac/baseline.json`;
outcomes PASS/CHANGED/FAIL, `--promote` after human review.
- docs/META-ACCOUNTS-API.md — separate API for accountscenter.meta.com
(linkage/security surface; credential-isolated from phone-OTP).
- docs/PHONE-OTP.md — phone-number OTP login flow for muse.ai (proven on bl).
- `bin/accounts-health.py` — per-account CDP session probe (runs inside the netns).
- `bin/accounts-health.sh` — aggregates account vitality from ACCOUNTS.md,
signs + POSTs to the board health ingest (systemd timer, every 15 min).
- `bin/muse` (or `box muse`) — native terminal entrypoint for muse-cli; supports global lookups (`muse status`, `muse threads`, `muse unread`, `muse lookup`, `muse passkey`, `muse tmux`), per-account REPL chats (`muse <account> chat`), and automated CDP cookie extraction.
- `bin/super-cli.py` (`box` or `super`) — unified orchestrator CLI for fleet health (`box fleet`), seamless lookups (`box lookup`), thread inspections (`box thread list/view`), background tmux (`box tmux`), and passkey reference (`box passkey`).
- `bin/muse-cli-node <node> [args]` — runs muse-cli inside node's netns with dedicated Cloudflare WARP egress & auto-refreshing cookies.
- `bin/refresh-node-cookies.py <node>` — extracts fresh cookies from running Chromium CDP in netns into `~/.config/muse-cli/<node>/cookies.txt`.
- `bin/muse_hybrid.py` — programmatic hybrid bridge combining fast gateway calls with CDP fallbacks.
- `bin/muse-tmux.py` (`box tmux` / `muse tmux`) — shared tmux socket manager (`/tmp/tmux-muse.sock`) for agent background execution, pipe-pane logging, hybrid netns/container execution, and session pruning.
- `bin/agent_md.py` — CLI & library for auditing, reading, writing, and synchronizing agent `.md` drive files (`SOUL.md`, `PROACTIVE_PREFERENCES.md`, `HEARTBEAT.md`, etc.) across containers via Hatch WebSocket RPC.
- `bin/agent-drive-watchdog.py` — background drive watchdog and auto-healing daemon (every 10m via `agent-drive-watchdog.timer`).
- `bin/swarm_worker/` & `bin/swarm-worker-supervise.sh` — supervised autonomous swarm worker daemon executing queued tasks in a hard sandbox.
- `bin/fleet-alert-relay.sh` — idempotent alert relay with 3-gate deduplication (watermark + 10m TTL hash + receipt verification) posting critical conditions to `#lobby`.
- `shared/operators/` — canonical operator drive markdown templates ensuring agents maintain autonomous loops, active supervision, and self-healing reflexes.
- docs/OPERATOR-DRIVE-RUNBOOK.md — operator runbook for auditing and modifying agent `.md` files via Hatch WebSocket RPC and SSH reverse tunnels.
- docs/BOX-WEB-SURFACE-GUIDE.md — Box web surface (`box.muse-dev.online`), passkey reality (single .txt file on VM), and agent approval flow.
- docs/HYBRID-GATEWAY-ADAPTATION.md — architectural guide on the muse-cli fast gateway adaptation and per-node egress isolation.
- docs/AGENT-TOOLING.md — guide to agent delegation, seamless lookups, shared tmux background tooling, Work Orders (`[WO:...]`), and prompt envelope execution.
- `bin/docs-lookup.py` (`box docs` / `super docs` / `docs-lookup`) — internal documentation, sentence structure grammar, regex passing engine, and assistive surfaces for `box.muse-dev.online`.
- `docs_internal/` — dual `.md` and structured `.json` lookup database for autonomous agents, protocol definitions, regex fixtures, and surface selectors.
## Verification checklist
-85
View File
@@ -1,85 +0,0 @@
#!/usr/bin/env python3
"""Per-account session vitality check (runs INSIDE the node's netns).
Usage: sudo ip netns exec warp-<node> python3 accounts-health.py <cdp_port>
Probes the account's browser via CDP:
- browser reachable
- muse.ai tab present
- login markers (heuristic: "Log in" button vs user content)
Prints JSON to stdout. Exit 0 on success, 1 if the browser is unreachable.
"""
import json, sys, urllib.request
CDP_PORT = int(sys.argv[1]) if len(sys.argv) > 1 else 9410
def http(path, timeout=5):
with urllib.request.urlopen(f"http://127.0.0.1:{CDP_PORT}{path}",
timeout=timeout) as r:
return json.loads(r.read())
result = {"browser_up": False, "muse_tab": False,
"session_alive": None, "title": "", "url": "", "detail": ""}
try:
http("/json/version")
result["browser_up"] = True
except Exception as e:
result["detail"] = f"CDP unreachable: {e}"
print(json.dumps(result)); sys.exit(1)
try:
tabs = http("/json/list")
except Exception as e:
result["detail"] = f"tab list failed: {e}"
print(json.dumps(result)); sys.exit(0)
muse_tab = None
for t in tabs:
url = t.get("url", "")
if "muse.ai" in url and t.get("type") == "page":
muse_tab = t
break
if not muse_tab:
result["detail"] = "no muse.ai tab open"
print(json.dumps(result)); sys.exit(0)
result["muse_tab"] = True
result["url"] = muse_tab.get("url", "")
result["title"] = muse_tab.get("title", "")
# DOM heuristic via the tab's debugger socket
try:
import websocket
ws = websocket.create_connection(muse_tab["webSocketDebuggerUrl"], timeout=10)
js = """JSON.stringify({
loginButtons: [...document.querySelectorAll('button')].filter(
b => /^\\s*log\\s*in\\s*$/i.test(b.innerText)).map(b => b.innerText.trim()),
hasAvatar: !!document.querySelector(
'img[alt*="avatar" i], [data-testid*="avatar" i], [aria-label*="profile" i]'),
title: document.title,
url: location.href
})"""
ws.send(json.dumps({"id": 1, "method": "Runtime.evaluate",
"params": {"expression": js, "returnByValue": True}}))
resp = json.loads(ws.recv())
ws.close()
dom = json.loads(resp["result"]["result"]["value"])
# Heuristic: login buttons present + no avatar => logged out.
# access/verification gate => needs verification.
# No login buttons (or avatar present) => likely logged in.
if "access/verification" in dom["url"]:
result["session_alive"] = False
result["detail"] = "age verification gate (access/verification)"
elif dom["loginButtons"] and not dom["hasAvatar"]:
result["session_alive"] = False
result["detail"] = f"login wall visible: {dom['loginButtons'][:2]}"
else:
result["session_alive"] = True
result["detail"] = "no login wall detected"
except Exception as e:
result["detail"] = f"DOM check failed: {e}"
print(json.dumps(result))
-97
View File
@@ -1,97 +0,0 @@
#!/usr/bin/env bash
# accounts-health.sh — account session vitality reporter for the front-door network.
#
# Reads ACCOUNTS.md, probes each account's browser via CDP (inside its NetVM
# netns) for session liveness, and emits a JSON report. Optionally signs and
# POSTs it to the board health ingest, following the health-report.sh
# convention: payload is <machine>\n<ts>\n<facts-json>, namespace "health".
#
# Usage: accounts-health.sh [--no-post]
#
# Cron (on bl, every 15 min):
# */15 * * * * ~/Projects/NetVM/bin/accounts-health.sh >/dev/null 2>&1
#
# Health key setup (once, on bl):
# ssh-keygen -t ed25519 -N "" -f ~/.ssh/muse-health
# # operator registers the pubkey on the VM:
# echo "bl $(cat ~/.ssh/muse-health.pub)" \
# | ssh super@34.139.37.135 "sudo tee -a /srv/board/health_signers"
set -u
NETVM_DIR="${NETVM_DIR:-$HOME/Projects/NetVM}"
ACCOUNTS="$NETVM_DIR/ACCOUNTS.md"
CHECKER="$NETVM_DIR/bin/accounts-health.py"
MACHINE="${MUSE_MACHINE:-bl}"
KEY="${HEALTH_KEY:-$HOME/.ssh/muse-health}"
ENDPOINT="${HEALTH_ENDPOINT:-https://board.muse-dev.online/api/health/report}"
POST=1
[ "${1:-}" = "--no-post" ] && POST=0
[ -f "$ACCOUNTS" ] || { echo "accounts-health: $ACCOUNTS missing" >&2; exit 1; }
[ -f "$CHECKER" ] || { echo "accounts-health: $CHECKER missing" >&2; exit 1; }
TS="$(date +%s)"
TMP="$(mktemp -d)"
trap 'rm -rf "$TMP"' EXIT
# Parse ACCOUNTS.md pipe table, probe each account inside its netns
NETVM_DIR="$NETVM_DIR" python3 - > "$TMP/facts.json" <<'PYEOF'
import json, os, subprocess, time
netvm = os.environ["NETVM_DIR"]
rows = []
for line in open(os.path.join(netvm, "ACCOUNTS.md")):
line = line.strip()
if not line.startswith("|"):
continue
cells = [c.strip() for c in line.strip("|").split("|")]
if len(cells) < 13 or cells[0] in ("agent", "-------", ""):
continue
if cells[10] not in ("", "-"):
rows.append({"agent": cells[0], "node": cells[1],
"status": cells[8], "cdp_port": cells[10]})
checker = os.path.join(netvm, "bin", "accounts-health.py")
out = {}
for a in rows:
try:
r = subprocess.run(
["sudo", "-n", "ip", "netns", "exec", f"warp-{a['node']}",
"python3", checker, a["cdp_port"]],
capture_output=True, text=True, timeout=60)
res = json.loads(r.stdout.strip().splitlines()[-1])
res["registry_status"] = a["status"]
out[a["agent"]] = res
except Exception as e:
out[a["agent"]] = {"browser_up": False, "session_alive": None,
"detail": f"probe failed: {e}",
"registry_status": a["status"]}
print(json.dumps({"accounts": out, "checked_at": int(time.time())}, indent=2))
PYEOF
if [ "$POST" -eq 0 ]; then
cat "$TMP/facts.json"
exit 0
fi
if [ ! -f "$KEY" ]; then
echo "accounts-health: $KEY missing — printing JSON, not posting (see header for key setup)" >&2
cat "$TMP/facts.json"
exit 0
fi
printf '%s\n%s\n' "$MACHINE" "$TS" > "$TMP/payload"
FACTS_JSON="$(cat "$TMP/facts.json")"
printf '%s' "$FACTS_JSON" >> "$TMP/payload"
ssh-keygen -Y sign -f "$KEY" -n health "$TMP/payload" >/dev/null 2>&1
SIG="$(cat "$TMP/payload.sig")"
python3 - "$MACHINE" "$TS" "$FACTS_JSON" "$SIG" <<'PYEOF' > "$TMP/body.json"
import json, sys
machine, ts, facts_json, sig = sys.argv[1], int(sys.argv[2]), sys.argv[3], sys.argv[4]
body = {"machine": machine, "ts": ts,
"facts": json.loads(facts_json), "facts_json": facts_json,
"signature": sig}
print(json.dumps(body))
PYEOF
curl -s -X POST "$ENDPOINT" -H 'Content-Type: application/json' \
--data @"$TMP/body.json" | head -c 300
echo
+654
View File
@@ -0,0 +1,654 @@
#!/usr/bin/env python3
"""agent-cognitive-probe.py — Real-time Cognitive Sensing and Agent Menu Navigation.
Directly probes Cloud Muse / Hatch browser runtime via CDP:
1. Passive Cognitive Sensing (zero-click):
- Token streaming / generation state (stop button presence)
- Typing / thinking indicators
- Status text displayed under/beside avatar (Connected, Thinking, Working)
- Active context (Main chat vs Side chats with thread titles & snippets)
- Input wait / parked approval detection
2. Active Profile Menu Navigation:
- Status panel sliding surface navigation
- Tabs: Activity (tasks & processes), Upcoming (timers & recurring cron loops),
Approvals, and Identity.
3. Cognitive Lock Gate:
- Protects single-threaded thought process from interruptions.
"""
import argparse
import json
import os
import subprocess
import sys
import time
import urllib.request
# Node to pinned CDP port mapping
NODE_CDP_PORTS = {
"muse": 9410,
"pip": 9420,
"646": 9430,
"opm": 9440,
"def": 9450,
"dev": 9455,
"muse-main": 9410,
}
REMOTE_HOST = "100.123.153.75" # bl control node
def is_running_on_bl():
"""Detect if we are running locally on bl or on tp/remote."""
try:
import socket
hn = socket.gethostname().lower()
if "bl" in hn:
return True
except Exception:
pass
# Check if network namespaces exist locally
return os.path.exists("/var/run/netns/warp-muse") or os.path.exists("/run/netns/warp-muse")
def run_cdp_eval_inside_netns(node, js_code, timeout=8):
"""Executes a JS snippet against the node's browser via CDP inside its netns."""
port = NODE_CDP_PORTS.get(node)
if not port:
return {"error": f"Unknown node '{node}'"}
# Python runner script to execute inside the target environment
runner_code = f'''
import json, sys, urllib.request
try:
import websocket
except ImportError:
print(json.dumps({{"error": "websocket package missing"}}))
sys.exit(1)
try:
with urllib.request.urlopen("http://127.0.0.1:{port}/json/list", timeout=3) as r:
targets = json.load(r)
pages = [t for t in targets if t.get("type") == "page"]
if not pages:
print(json.dumps({{"error": "No active page target"}}))
sys.exit(0)
ws_url = pages[0]["webSocketDebuggerUrl"]
ws = websocket.create_connection(ws_url, timeout={timeout})
ws.send(json.dumps({{
"id": 1,
"method": "Runtime.evaluate",
"params": {{
"expression": {json.dumps(js_code)},
"returnByValue": True,
"awaitPromise": True
}}
}}))
res = None
for _ in range(30):
msg = json.loads(ws.recv())
if msg.get("id") == 1:
res = msg.get("result", {{}}).get("result", {{}}).get("value")
break
print(json.dumps({{"ok": True, "value": res}}))
except Exception as e:
print(json.dumps({{"error": str(e)}}))
'''
if is_running_on_bl():
cmd = ["sudo", "ip", "netns", "exec", f"warp-{node}", "python3", "-c", runner_code]
else:
# Wrap via ssh to bl
# Use python3 on bl directly executing inside netns
remote_cmd = f"sudo ip netns exec warp-{node} python3 -c {subprocess.list2cmdline([runner_code])}"
cmd = ["ssh", "-q", f"super@{REMOTE_HOST}", remote_cmd]
try:
proc = subprocess.run(cmd, capture_output=True, text=True, timeout=timeout + 5)
if proc.returncode != 0 and not proc.stdout:
return {"error": proc.stderr.strip() or f"Process exited with {proc.returncode}"}
# Parse output line that contains valid json
for line in proc.stdout.strip().splitlines():
line = line.strip()
if line.startswith("{") and line.endswith("}"):
try:
data = json.loads(line)
if "ok" in data:
return data["value"]
if "error" in data:
return {"error": data["error"]}
except Exception:
continue
return {"error": proc.stdout.strip() or proc.stderr.strip()}
except subprocess.TimeoutExpired:
return {"error": "CDP probe timed out"}
except Exception as e:
return {"error": str(e)}
JS_PASSIVE_COGNITIVE = """(() => {
// 1. Generation & Thinking signals
const stopBtn = document.querySelector('[data-testid="hatch-composer-stop-button"]');
const isGenerating = !!stopBtn;
const typingEl = document.querySelector('[data-testid="hatch-chat-typing-indicator"]');
const isTyping = !!typingEl && typingEl.textContent.trim().length > 0;
// 2. Avatar / Status text
const statusTextEl = document.querySelector('.group\\\\/status-avatar span, [class*="status-avatar"] span, span[class*="text-body-status"]');
const avatarStatus = statusTextEl ? (statusTextEl.innerText || '').trim() : '';
// 3. Thread context & active URL
const url = window.location.href;
const isMainChat = url === 'https://muse.ai/' || url.endsWith('/thread/new');
const pageTitle = document.title;
// 4. Side chats overview
const sideRows = Array.from(document.querySelectorAll('[data-testid="hatch-thread-row"]')).map(r => {
const text = (r.innerText || '').trim().replace(/\\\\n+/g, ' — ');
return text;
}).slice(0, 5);
// 5. Input wait / Parked prompt cards
// Detect buttons asking for approval / input in the message flow
const actionBtns = Array.from(document.querySelectorAll('div[data-message-id] button')).map(b => (b.innerText || '').trim()).filter(t => /approve|confirm|proceed|resume|start|allow/i.test(t));
const isInputWait = actionBtns.length > 0 || /asking for input|pending approval/i.test(document.body.innerText.slice(-600));
// 6. Status panel state
const panel = document.querySelector('[data-testid="hatch-status-panel-sliding-surface"]');
const panelOpen = !!panel && panel.getBoundingClientRect().width > 0;
return {
url: url,
title: pageTitle,
is_main_chat: isMainChat,
is_generating: isGenerating,
is_typing: isTyping,
avatar_status: avatarStatus,
is_input_wait: isInputWait,
pending_actions: actionBtns,
panel_open: panelOpen,
side_chats: sideRows
};
})()"""
def get_passive_cognitive_state(node):
"""Gathers passive cognitive signals without altering UI state."""
if not is_running_on_bl():
try:
cmd = ["ssh", "-q", "-o", "ConnectTimeout=5", f"super@{REMOTE_HOST}",
f"python3 /home/super/Projects/NetVM/bin/agent-cognitive-probe.py status {node} --json"]
proc = subprocess.run(cmd, capture_output=True, text=True, timeout=10)
if proc.returncode == 0 and proc.stdout.strip():
data = json.loads(proc.stdout.strip())
if isinstance(data, list) and len(data) > 0:
return data[0]
elif isinstance(data, dict):
return data
except Exception as e:
return {
"node": node,
"status": "DARK",
"error": f"Remote delegation failed: {e}",
"cognitive_lock": False,
"lock_reason": None,
}
raw = run_cdp_eval_inside_netns(node, JS_PASSIVE_COGNITIVE)
if not isinstance(raw, dict) or "error" in raw:
return {
"node": node,
"status": "DARK",
"error": raw.get("error", "Unknown error") if isinstance(raw, dict) else str(raw),
"cognitive_lock": False,
"lock_reason": None,
}
is_generating = raw.get("is_generating", False)
is_typing = raw.get("is_typing", False)
is_input_wait = raw.get("is_input_wait", False)
avatar_status = raw.get("avatar_status", "")
is_main = raw.get("is_main_chat", True)
# Determine synthesized cognitive state
if is_generating or is_typing or "thinking" in avatar_status.lower():
state = "THINKING"
locked = True
reason = "Agent is actively generating tokens / thinking (stop button active)"
elif is_input_wait:
state = "INPUT_WAIT"
locked = True
reason = "Agent is waiting for operator or system input on a parked prompt"
elif "working" in avatar_status.lower() or "making" in avatar_status.lower():
state = "WORKING"
locked = True
reason = f"Avatar status indicates work in progress: '{avatar_status}'"
elif not is_main:
state = "SIDECHAT_IDLE"
locked = False
reason = None
else:
state = "IDLE"
locked = False
reason = None
return {
"node": node,
"status": state,
"cognitive_lock": locked,
"lock_reason": reason,
"details": raw,
}
JS_NAVIGATE_MENU_TEMPLATE = """(async () => {
// 1. Ensure panel is open
let panel = document.querySelector('[data-testid="hatch-status-panel-sliding-surface"]');
if (!panel) {
// Try clicking avatar or trigger
const avatarBtn = document.querySelector('[role="img"][aria-label*="avatar"], button[aria-label*="avatar"], .group\\\\/status-avatar');
if (avatarBtn) avatarBtn.click();
await new Promise(r => setTimeout(r, 350));
}
// 2. Click requested tab if specified
const targetTab = "%(tab)s";
if (targetTab && targetTab !== "all") {
const btn = document.querySelector('button[aria-label="' + targetTab + '"]');
if (btn) {
btn.click();
await new Promise(r => setTimeout(r, 400));
}
}
panel = document.querySelector('[data-testid="hatch-status-panel-sliding-surface"]');
if (!panel) return {error: "Status panel not rendered"};
// Read full text and structural items
const rawText = panel.innerText || '';
const lines = rawText.split('\\n').map(s => s.trim()).filter(Boolean);
// Extract activity / timer items
const items = [];
const buttons = Array.from(panel.querySelectorAll('button')).filter(b => !b.getAttribute('aria-label') && (b.innerText || '').length > 0);
buttons.forEach(b => {
const text = (b.innerText || '').trim();
const parts = text.split('\\n').map(s => s.trim()).filter(Boolean);
if (parts.length >= 2) {
items.push({
title: parts[0],
detail: parts[1],
time: parts.length > 2 ? parts[2] : null
});
}
});
return {
raw_text: rawText,
lines: lines,
items: items
};
})()"""
def get_agent_menu(node, tab="all"):
"""Navigates and extracts data from the agent profile / status panel."""
if not is_running_on_bl():
try:
cmd = ["ssh", "-q", "-o", "ConnectTimeout=5", f"super@{REMOTE_HOST}",
f"python3 /home/super/Projects/NetVM/bin/agent-cognitive-probe.py menu {node} {tab} --json"]
proc = subprocess.run(cmd, capture_output=True, text=True, timeout=15)
if proc.returncode == 0 and proc.stdout.strip():
return json.loads(proc.stdout.strip())
except Exception as e:
return {
"node": node,
"tab": tab,
"data": {"error": f"Remote delegation failed: {e}"}
}
tab_map = {
"activity": "Activity",
"upcoming": "Upcoming",
"approvals": "Approvals",
"identity": "Identity",
"all": "all",
}
target_tab = tab_map.get(tab.lower(), "Activity")
# If all is requested, gather activity, upcoming, and approvals
if target_tab == "all":
result = {}
for sub_tab in ["Activity", "Upcoming", "Approvals"]:
js = JS_NAVIGATE_MENU_TEMPLATE % {"tab": sub_tab}
tab_res = run_cdp_eval_inside_netns(node, js)
result[sub_tab.lower()] = tab_res
return {
"node": node,
"menu": result
}
js = JS_NAVIGATE_MENU_TEMPLATE % {"tab": target_tab}
res = run_cdp_eval_inside_netns(node, js)
return {
"node": node,
"tab": target_tab,
"data": res
}
JS_LIVE_SCREEN = """(() => {
const ps = Array.from(document.querySelectorAll('p')).map(p => (p.innerText || '').trim()).filter(Boolean);
const actionBtns = Array.from(document.querySelectorAll('div[data-message-id] button, [role="log"] button, div[role="status"] button')).map(b => (b.innerText || '').trim()).filter(Boolean);
const stopBtn = !!document.querySelector('[data-testid="hatch-composer-stop-button"]');
const typing = !!document.querySelector('[data-testid="hatch-chat-typing-indicator"]:not(:empty)');
const statusTextEl = document.querySelector('.group\\\\/status-avatar span, [class*="status-avatar"] span, span[class*="text-body-status"]');
const avatarStatus = statusTextEl ? (statusTextEl.innerText || '').trim() : '';
return {
title: document.title,
url: window.location.href,
is_main_chat: window.location.href === 'https://muse.ai/' || window.location.href.endsWith('/thread/new'),
is_generating: stopBtn,
is_typing: typing,
avatar_status: avatarStatus,
recent_paragraphs: ps.slice(-8),
action_buttons: actionBtns.filter(t => /approve|confirm|proceed|resume|start|allow|review/i.test(t))
};
})()"""
def get_agent_live_screen(node: str) -> dict:
"""Extracts live active chat text, thoughts, prompt blocks, and threads."""
if not is_running_on_bl():
try:
cmd = ["ssh", "-q", "-o", "ConnectTimeout=5", f"super@{REMOTE_HOST}",
f"python3 /home/super/Projects/NetVM/bin/agent-cognitive-probe.py read {node} --json"]
proc = subprocess.run(cmd, capture_output=True, text=True, timeout=15)
if proc.returncode == 0 and proc.stdout.strip():
return json.loads(proc.stdout.strip())
except Exception as e:
return {"node": node, "error": f"Remote delegation failed: {e}"}
screen = run_cdp_eval_inside_netns(node, JS_LIVE_SCREEN)
if not isinstance(screen, dict) or "error" in screen:
return {"node": node, "error": screen.get("error", "Failed to inspect screen") if isinstance(screen, dict) else str(screen)}
sidechats = []
try:
cli_cmd = ["/home/super/Projects/NetVM/bin/muse-cli-node", node, "threads"]
proc = subprocess.run(cli_cmd, capture_output=True, text=True, timeout=8)
if proc.returncode == 0 and proc.stdout.strip():
threads_data = json.loads(proc.stdout.strip())
if isinstance(threads_data, list):
sidechats = threads_data[:8]
except Exception:
pass
return {
"node": node,
"screen": screen,
"sidechats": sidechats
}
JS_NAVIGATE_MAIN_CHAT = """(() => {
try {
const url = window.location.href;
if (url === 'https://muse.ai/' || url === 'https://muse.ai/thread/new') {
return { ok: true, already_main: true };
}
const candidates = Array.from(document.querySelectorAll('a, button, [role="button"], [data-testid]'));
const mainBtn = candidates.find(b => {
const t = (b.innerText || '').trim().toLowerCase();
return t === 'main chat' || b.getAttribute('aria-label') === 'Main chat' || b.getAttribute('data-testid') === 'hatch-sidebar-main-chat';
});
if (mainBtn) {
mainBtn.click();
return { ok: true, method: 'click' };
}
window.location.href = 'https://muse.ai/';
return { ok: true, method: 'navigate' };
} catch (e) {
return { error: String(e) };
}
})()"""
def navigate_to_main_chat(node: str) -> dict:
"""Navigates the agent browser session back to Main Chat (https://muse.ai/)."""
if not is_running_on_bl():
try:
cmd = ["ssh", "-q", "-o", "ConnectTimeout=5", f"super@{REMOTE_HOST}",
f"python3 /home/super/Projects/NetVM/bin/agent-cognitive-probe.py nav-main {node} --json"]
proc = subprocess.run(cmd, capture_output=True, text=True, timeout=10)
if proc.returncode == 0 and proc.stdout.strip():
return json.loads(proc.stdout.strip())
except Exception as e:
return {"error": str(e)}
return run_cdp_eval_inside_netns(node, JS_NAVIGATE_MAIN_CHAT)
def cmd_nav_main(args):
node = getattr(args, "agent", None) or getattr(args, "node", None)
res = navigate_to_main_chat(node)
if getattr(args, "json", False):
print(json.dumps(res, indent=2))
else:
if isinstance(res, dict) and res.get("ok"):
m = res.get("method") or ("already in main chat" if res.get("already_main") else "default")
print(f"✓ Refocused agent '{node}' to Main Chat ({m})")
else:
err = res.get("error") if isinstance(res, dict) else str(res)
print(f"❌ Failed to refocus '{node}' to Main Chat: {err}")
def cmd_read(args):
node = getattr(args, "agent", None) or getattr(args, "node", None)
data = get_agent_live_screen(node)
if getattr(args, "json", False):
print(json.dumps(data, indent=2))
return
if "error" in data:
print(f"\n❌ Error inspecting screen for {node}: {data['error']}\n")
return
screen = data.get("screen", {})
sidechats = data.get("sidechats", [])
url = screen.get("url", "")
mode = "Main Chat" if screen.get("is_main_chat") else "Side Chat"
title = screen.get("title", "")
avatar = screen.get("avatar_status") or "Connected"
gen = "🧠 GENERATING" if screen.get("is_generating") else ("💭 TYPING" if screen.get("is_typing") else "🟢 SETTLED")
print(f"\n=== LIVE THOUGHT STREAM & ACTIVE CHAT: {node.upper()} ===")
print(f"Context: {mode} ({url})")
print(f"Title: {title}")
print(f"Status: {avatar} | State: {gen}")
paragraphs = screen.get("recent_paragraphs", [])
if paragraphs:
print("\n--- ACTIVE CONVERSATION & THOUGHT PARAGRAPHS ---")
for p in paragraphs:
print(f" • {p}\n")
else:
print("\n (No text paragraphs visible in current viewport)")
actions = screen.get("action_buttons", [])
if actions:
print("--- PENDING ACTION CARDS / APPROVAL BUTTONS ---")
for a in actions:
print(f" ⚠️ [PROMPT ACTION] {a}")
print()
if sidechats:
print("--- RECENT SIDE CHATS & TOPICS ---")
for sc in sidechats:
sid = sc.get("session_id", "")[:8]
stitle = sc.get("title") or "(Untitled sidechat)"
upd = sc.get("updated", "")
print(f" • [{sid}] {stitle} ({upd})")
print()
def format_cognitive_badge(status):
badges = {
"IDLE": "🟢 IDLE",
"SIDECHAT_IDLE": "💬 SIDE_IDLE",
"THINKING": "🧠 THINKING",
"INPUT_WAIT": "⏸️ INPUT_WAIT",
"BUSY_SIDECHAT": "💬 SIDECHAT",
"WORKING": "⚙️ WORKING",
"DARK": "⚫ DARK",
}
return badges.get(status, f"❓ {status}")
def cmd_status(args):
nodes = getattr(args, "agents", None) or getattr(args, "nodes", None) or ["muse", "pip", "646", "opm", "dev", "def"]
results = []
for n in nodes:
results.append(get_passive_cognitive_state(n))
if getattr(args, "json", False):
print(json.dumps(results, indent=2))
return
print("\n=== AGENT COGNITIVE SENSOR (LIVE DOM PROBE) ===")
print(f"{'AGENT':<10} {'COGNITIVE STATE':<18} {'LOCK':<8} {'UNDER-AVATAR':<15} {'DETAILS / CURRENT THOUGHT':<40}")
print("-" * 95)
for r in results:
node = r["node"]
status = r["status"]
badge = format_cognitive_badge(status)
locked = "LOCKED" if r.get("cognitive_lock") else "OPEN"
det = r.get("details", {})
avatar_status = det.get("avatar_status", "-")
reason = r.get("lock_reason") or det.get("title", "Settled")
if status == "DARK":
reason = r.get("error", "CDP unreachable")
avatar_status = "DARK"
print(f"{node:<10} {badge:<18} {locked:<8} {avatar_status:<15} {reason[:40]:<40}")
print()
def cmd_menu(args):
node = getattr(args, "agent", None) or getattr(args, "node", None)
tab = getattr(args, "tab", "activity") or "activity"
data = get_agent_menu(node, tab)
if getattr(args, "json", False):
print(json.dumps(data, indent=2))
return
print(f"\n=== AGENT MENU: {node.upper()} (TAB: {tab.upper()}) ===")
if tab.lower() == "all":
menu = data.get("menu", {})
for tname, tdata in menu.items():
print(f"\n--- {tname.upper()} ---")
if isinstance(tdata, dict) and "error" in tdata:
print(f" Error: {tdata['error']}")
elif isinstance(tdata, dict):
items = tdata.get("items", [])
if items:
for it in items:
t_str = f" [{it['time']}]" if it.get("time") else ""
print(f" • {it['title']}: {it['detail']}{t_str}")
else:
raw = tdata.get("raw_text", "")
for line in raw.split("\n"):
if line.strip():
print(f" {line.strip()}")
print()
return
tdata = data.get("data", {})
if isinstance(tdata, dict) and "error" in tdata:
print(f" Error: {tdata['error']}")
elif isinstance(tdata, dict):
items = tdata.get("items", [])
if items:
for it in items:
t_str = f" [{it['time']}]" if it.get("time") else ""
print(f" • {it['title']}: {it['detail']}{t_str}")
else:
raw = tdata.get("raw_text", "")
for line in raw.split("\n"):
if line.strip():
print(f" {line.strip()}")
print()
def cmd_lock(args):
node = getattr(args, "agent", None) or getattr(args, "node", None)
state = get_passive_cognitive_state(node)
if args.json:
print(json.dumps(state, indent=2))
else:
if state.get("cognitive_lock"):
print(f"🔴 COGNITIVE LOCK ENGAGED on '{node}' ({state['status']}): {state.get('lock_reason')}")
else:
print(f"🟢 COGNITIVELY IDLE: Agent '{node}' is free to receive new work without interruption.")
if state.get("cognitive_lock") and not getattr(args, "force", False):
sys.exit(1)
sys.exit(0)
def main():
parser = argparse.ArgumentParser(description="Real-time Cognitive Sensing and Agent Menu Navigation")
sub = parser.add_subparsers(dest="command")
p_status = sub.add_parser("status", help="Show cognitive state for all or selected agents")
p_status.add_argument("nodes", nargs="*", help="Optional agent names")
p_status.add_argument("--json", action="store_true", help="Output JSON")
p_menu = sub.add_parser("menu", help="Navigate agent profile menu (tasks, timers, approvals, identity)")
p_menu.add_argument("node", help="Agent name (muse, pip, 646, opm, dev, def)")
p_menu.add_argument("tab", nargs="?", default="activity", choices=["activity", "upcoming", "approvals", "identity", "all"], help="Menu tab to view")
p_menu.add_argument("--json", action="store_true", help="Output JSON")
p_lock = sub.add_parser("lock", help="Check cognitive lock before dispatching work")
p_lock.add_argument("node", help="Agent name")
p_lock.add_argument("--force", action="store_true", help="Bypass lock check")
p_lock.add_argument("--json", action="store_true", help="Output JSON")
p_read = sub.add_parser("read", help="Extract live active chat text, thought stream, and side chats")
p_read.add_argument("node", help="Agent name")
p_read.add_argument("--json", action="store_true", help="Output JSON")
p_nav = sub.add_parser("nav-main", help="Navigate agent browser session back to Main Chat")
p_nav.add_argument("node", help="Agent name")
p_nav.add_argument("--json", action="store_true", help="Output JSON")
args = parser.parse_args()
if not args.command:
# Default to status
args.nodes = []
args.json = False
cmd_status(args)
return
if args.command == "status":
cmd_status(args)
elif args.command == "menu":
cmd_menu(args)
elif args.command == "read":
cmd_read(args)
elif args.command == "lock":
cmd_lock(args)
elif args.command == "nav-main":
cmd_nav_main(args)
if __name__ == "__main__":
main()
-200
View File
@@ -1,200 +0,0 @@
#!/usr/bin/env python3
"""
agent-drive-watchdog.py — Automated Drive Watchdog & Healing Daemon for Muse Agents.
Monitors operational DRIVE scores across the agent fleet (muse, pip, 646, opm, def, dev)
via Hatch WebSocket RPC. If any agent's DRIVE score drops below 100 or critical drive
files (HEARTBEAT.md, PROACTIVE_PREFERENCES.md, SOUL.md) are degraded or missing:
1. Detects degraded state and missing checklist/preferences.
2. Selectively auto-heals core drive files using canonical shared templates.
3. Preserves MEMORY.md and agent-generated workspace files.
4. Records state & healing history to /tmp/agent-drive-watchdog.json.
5. Emits structured telemetry to stdout/journal.
Can be run:
- Once: python3 bin/agent-drive-watchdog.py --once
- Continuous loop: python3 bin/agent-drive-watchdog.py --interval 600
- Via systemd timer: agent-drive-watchdog.timer (every 10m)
"""
import argparse
import datetime
import json
import logging
import os
import sys
import time
from pathlib import Path
# Add NetVM bin to path
BASE_DIR = Path(__file__).resolve().parent.parent
sys.path.insert(0, str(BASE_DIR / "bin"))
from agent_md import audit_agents, write_md, read_md, SHARED_OPERATORS, VALID_ACCOUNTS
STATE_FILE = Path("/tmp/agent-drive-watchdog.json")
LOG_FILE = Path("/tmp/agent-drive-watchdog.log")
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s [%(levelname)s] %(message)s",
handlers=[
logging.StreamHandler(sys.stdout),
logging.FileHandler(LOG_FILE, mode="a", encoding="utf-8")
]
)
def load_state() -> dict:
if STATE_FILE.exists():
try:
return json.loads(STATE_FILE.read_text(encoding="utf-8"))
except Exception as e:
logging.warning(f"Failed to read existing state file: {e}")
return {
"last_run": None,
"history": [],
"agent_status": {},
}
def save_state(state: dict):
try:
# Keep last 50 history entries
if len(state.get("history", [])) > 50:
state["history"] = state["history"][-50:]
STATE_FILE.write_text(json.dumps(state, indent=2), encoding="utf-8")
except Exception as e:
logging.error(f"Failed to save state file: {e}")
def heal_agent_drive(account: str, audit_data: dict) -> dict:
"""Selectively auto-heal core drive files for an agent."""
healed = []
issues = audit_data.get("issues", [])
# Check what specifically needs healing
need_soul = any("SOUL.md" in iss for iss in issues) or audit_data.get("files", {}).get("SOUL.md", {}).get("size", 0) < 1000
need_pro = any("PROACTIVE_PREFERENCES.md" in iss for iss in issues) or audit_data.get("files", {}).get("PROACTIVE_PREFERENCES.md", {}).get("size", 0) < 800
need_hb = any("HEARTBEAT.md" in iss for iss in issues) or audit_data.get("files", {}).get("HEARTBEAT.md", {}).get("size", 0) < 300
need_context = any("TOOLS.md or USER.md" in iss for iss in issues)
logging.info(f"[{account.upper()}] Auto-healing drive... need_soul={need_soul}, need_pro={need_pro}, need_hb={need_hb}, need_context={need_context}")
if need_soul:
soul_text = (SHARED_OPERATORS / "SOUL.md").read_text(encoding="utf-8")
res = write_md(account, "SOUL.md", soul_text, overwrite=True)
healed.append({"file": "SOUL.md", "bytes": res.get("bytes_written")})
if need_pro:
pro_text = (SHARED_OPERATORS / "PROACTIVE_PREFERENCES.md").read_text(encoding="utf-8")
res = write_md(account, "PROACTIVE_PREFERENCES.md", pro_text, overwrite=True)
healed.append({"file": "PROACTIVE_PREFERENCES.md", "bytes": res.get("bytes_written")})
if need_hb:
hb_text = (SHARED_OPERATORS / "HEARTBEAT.md").read_text(encoding="utf-8")
res = write_md(account, "HEARTBEAT.md", hb_text, overwrite=True)
healed.append({"file": "HEARTBEAT.md", "bytes": res.get("bytes_written")})
if need_context:
tools_text = (SHARED_OPERATORS / "TOOLS.md").read_text(encoding="utf-8")
res_t = write_md(account, "TOOLS.md", tools_text, overwrite=True)
healed.append({"file": "TOOLS.md", "bytes": res_t.get("bytes_written")})
user_text = (SHARED_OPERATORS / "USER.md").read_text(encoding="utf-8")
res_u = write_md(account, "USER.md", user_text, overwrite=True)
healed.append({"file": "USER.md", "bytes": res_u.get("bytes_written")})
return {"ok": True, "account": account, "healed": healed}
def run_cycle(auto_heal: bool = True) -> dict:
"""Run an audit and healing cycle across all fleet accounts."""
now_iso = datetime.datetime.now(datetime.timezone.utc).isoformat()
logging.info("Starting fleet drive watchdog audit cycle...")
state = load_state()
state["last_run"] = now_iso
try:
audit_results = audit_agents()
except Exception as e:
logging.error(f"Audit failed during cycle: {e}")
return {"ok": False, "error": str(e)}
cycle_report = {
"timestamp": now_iso,
"total_agents": len(audit_results),
"high_drive": 0,
"degraded": 0,
"healed_agents": [],
}
for account, a_data in audit_results.items():
score = a_data.get("drive_score", 0)
issues = a_data.get("issues", [])
status = "HIGH_DRIVE" if score == 100 else "DEGRADED"
if status == "HIGH_DRIVE":
cycle_report["high_drive"] += 1
logging.info(f"Agent {account.upper():6}: DRIVE score 100/100 (HIGH DRIVE)")
else:
cycle_report["degraded"] += 1
logging.warning(f"Agent {account.upper():6}: DRIVE score {score}/100 ({status}) - Issues: {', '.join(issues)}")
if auto_heal:
try:
heal_res = heal_agent_drive(account, a_data)
healed_files = [h["file"] for h in heal_res.get("healed", [])]
logging.info(f"Agent {account.upper():6}: Successfully healed files: {', '.join(healed_files)}")
cycle_report["healed_agents"].append({
"account": account,
"prior_score": score,
"healed_files": healed_files,
})
except Exception as e:
logging.error(f"Agent {account.upper():6}: Healing failed: {e}")
state["agent_status"][account] = {
"score": score,
"status": status,
"issues": issues,
"last_checked": now_iso,
}
state["history"].append(cycle_report)
save_state(state)
logging.info(f"Watchdog cycle complete. High drive: {cycle_report['high_drive']}/{cycle_report['total_agents']}. Degraded: {cycle_report['degraded']}. Healed: {len(cycle_report['healed_agents'])}.")
return cycle_report
def main():
parser = argparse.ArgumentParser(description="NetVM Automated Agent Drive Watchdog & Healing Daemon")
parser.add_argument("--once", action="store_true", help="Run a single audit/healing pass and exit")
parser.add_argument("--no-heal", action="store_true", help="Audit only; do not auto-heal degraded agents")
parser.add_argument("--interval", type=int, default=600, help="Loop interval in seconds (default: 600s / 10m)")
parser.add_argument("--status", action="store_true", help="Print recent watchdog status and exit")
args = parser.parse_args()
if args.status:
state = load_state()
print(json.dumps(state, indent=2))
return
if args.once:
run_cycle(auto_heal=not args.no_heal)
return
logging.info(f"Starting NetVM Agent Drive Watchdog daemon (interval: {args.interval}s)...")
while True:
try:
run_cycle(auto_heal=not args.no_heal)
except Exception as e:
logging.error(f"Unexpected error in watchdog loop: {e}", exc_info=True)
time.sleep(args.interval)
if __name__ == "__main__":
main()
-240
View File
@@ -1,240 +0,0 @@
#!/bin/bash
# GOLDEN PATH: container -> VM (34.139.37.135) -> bl (100.123.153.75) -> netns -> browser -> agent
# This script is the operator's heartbeat. If it stops, agents go dark.
# When debugging: trace each hop. Don't assume -- verify.
# Operator health monitor for Muse agents on bl.
# Checks each agent via API every 5 minutes. If unresponsive:
# 1. Restart the browser
# 2. Re-check
# 3. Log failure if still down
#
# Run via systemd timer or cron: */5 * * * * /home/super/Projects/NetVM/bin/agent-health.sh
#
# Agents are defined in ACCOUNTS.md. This script reads the active ones.
NETVM_BIN="$(cd "$(dirname "$0")" && pwd)"
LOG="/tmp/agent-health.log"
# 2026-10-04: per-node consecutive-API-failure counters. A single
# muse-chat-api.py failure must not kill a healthy browser (observed
# 2026-10-04 20:48:44 UTC: 646's browser killed on an API timeout while its
# CDP port was still listening). Require 2 CONSECUTIVE API failures before
# the kill path; the counter resets on any successful check.
STATE_DIR="/tmp/agent-health-state"
mkdir -p "$STATE_DIR" 2>/dev/null
check_cdp_port() {
# Verify CDP port is actually listening in the netns.
# A browser can be running but not bound to CDP (zombie state).
local node=$1
local cdp_port=$2
if sudo ip netns exec "warp-$node" ss -tln 2>/dev/null | grep -q ":$cdp_port "; then
return 0
else
return 1
fi
}
# Return codes: 0 = healthy, 1 = API check failed (port OK), 2 = CDP port down.
check_agent() {
local agent=$1
local node=$2
local cdp_port=$3
# First: verify CDP port is listening (catches zombie browsers)
if ! check_cdp_port "$node" "$cdp_port"; then
echo "$(date -Iseconds) $agent: FAIL (cdp port $cdp_port not listening)" >> "$LOG"
return 2
fi
# Then: try API messages command (lightweight check)
if timeout 30 "$NETVM_BIN/netvm-exec.sh" "$node" -- python3 "$NETVM_BIN/muse-chat-api.py" --account "$agent" messages 1 > /dev/null 2>&1; then
echo "$(date -Iseconds) $agent: OK" >> "$LOG"
return 0
else
echo "$(date -Iseconds) $agent: FAIL (api timeout)" >> "$LOG"
return 1
fi
}
# 2026-10-04: recent-relaunch guard. chromebox-watchdog.sh (2-min timer) and
# this script (5-min timer) could otherwise kill each other's fresh browsers:
# a browser just relaunched by the watchdog is still starting when this
# script's API check times out on it. Skip the kill path when the main
# browser process for this profile launched <2 min ago. Mirrors
# chromebox-watchdog.sh's relaunch-loop guard idiom (main process only:
# --remote-debugging-port present, no --type= flag; renderer/gpu children
# start later than the main process and must not satisfy this check).
recent_relaunch() {
local cdp_port=$1
local _pid _start _now
for _pid in $(pgrep -f "chromium.*--remote-debugging-port=${cdp_port}([[:space:]]|$)" 2>/dev/null); do
# Skip child processes (renderer, gpu, etc.) — only the main browser counts
if ps -o args= -p "$_pid" 2>/dev/null | grep -q -- "--type="; then
continue
fi
_start=$(date -d "$(ps -o lstart= -p "$_pid" 2>/dev/null)" +%s 2>/dev/null || echo 0)
_now=$(date +%s)
if [ $(( _now - _start )) -lt 120 ] && [ "$_start" -gt 0 ]; then
return 0
fi
done
return 1
}
restart_browser() {
local agent=$1
local cdp_port=$2
echo "$(date -Iseconds) $agent: restarting browser..." >> "$LOG"
# Kill existing by exact PIDs (pkill patterns are unreliable)
# Find chromium processes for this profile
for pid in $(pgrep -f "chromium.*profiles/$agent" 2>/dev/null); do
kill -9 "$pid" 2>/dev/null
done
sleep 3
# Verify port is free before restart
if sudo ip netns exec "warp-$agent" ss -tln 2>/dev/null | grep -q ":$cdp_port "; then
echo "$(date -Iseconds) $agent: WARNING - port $cdp_port still bound after kill" >> "$LOG"
fi
# Restart via netvm-chrome.sh in its own systemd scope.
# This oneshot service runs with KillMode=control-group, so anything
# spawned directly under it (nohup AND setsid both stay in the cgroup)
# is SIGKILLed when the service exits — observed 2026-10-03: every
# restart "recovered" then died at service teardown, looping forever.
# A transient scope escapes the service cgroup and survives.
# NOTE: systemd-run --scope WAITS for the scope's processes (even with
# --no-block, verified 2026-10-03), so background it — the scope itself
# is an independent unit and outlives the wrapper.
cd "$NETVM_BIN"
systemd-run --user --scope --unit="netvm-chrome-$agent-$(date +%s)" \
./netvm-chrome.sh --headless --cdp-port "$cdp_port" "$agent" https://muse.ai \
> "/tmp/bl-$agent.log" 2>&1 &
sleep 15
echo "$(date -Iseconds) $agent: browser restarted" >> "$LOG"
}
# 2026-10-06: restart circuit breaker. A restart that leaves the agent
# still failing is FUTILE (observed 2026-10-06: def's API check failed
# 155x while its CDP port was up; 57 kill+restart cycles murdered a
# healthy browser for an account-layer failure restarts cannot fix).
# After FUTILE_THRESHOLD consecutive futile restarts, open the circuit:
# stop killing/restarting and alert, until CIRCUIT_COOLDOWN seconds pass
# (one half-open probe restart) or any check succeeds. Manual reset:
# rm /tmp/agent-health-state/circuit-<agent> /tmp/agent-health-state/futile-<agent>
FUTILE_THRESHOLD=3
CIRCUIT_COOLDOWN=1800
# circuit_allows <agent>: return 0 if a restart may proceed now.
circuit_allows() {
local agent=$1 now opened retry_in
local cf="$STATE_DIR/circuit-$agent"
[ -f "$cf" ] || return 0
opened=$(cat "$cf" 2>/dev/null || echo 0)
now=$(date +%s)
if [ $(( now - opened )) -ge $CIRCUIT_COOLDOWN ]; then
echo "$(date -Iseconds) $agent: circuit half-open after ${CIRCUIT_COOLDOWN}s cooldown, one probe restart" >> "$LOG"
return 0
fi
retry_in=$(( (opened + CIRCUIT_COOLDOWN - now + 59) / 60 ))
echo "$(date -Iseconds) $agent: CIRCUIT OPEN - skipping kill/restart (restarts futile, probe retry in ~${retry_in}m; manual reset: rm $cf)" >> "$LOG"
return 1
}
# circuit_note_restart <agent> <ok|fail>: record a restart outcome.
circuit_note_restart() {
local agent=$1 outcome=$2 count=0
local ff="$STATE_DIR/futile-$agent" cf="$STATE_DIR/circuit-$agent"
if [ "$outcome" = "ok" ]; then
rm -f "$ff" "$cf" 2>/dev/null
return 0
fi
[ -f "$ff" ] && count=$(cat "$ff" 2>/dev/null || echo 0)
count=$(( count + 1 ))
echo "$count" > "$ff"
if [ "$count" -ge "$FUTILE_THRESHOLD" ]; then
date +%s > "$cf"
local msg="$agent: ALERT - $count consecutive futile restarts, circuit OPEN for ${CIRCUIT_COOLDOWN}s (no more kills until probe; manual reset: rm $cf $ff)"
echo "$(date -Iseconds) $msg" >> "$LOG"
echo "agent-health ALERT: $msg"
fi
}
# Allow sourcing for tests without running checks.
if [ "${AGENT_HEALTH_LIB_ONLY:-}" = "1" ]; then
return 0 2>/dev/null || exit 0
fi
# Main
echo "=== Health check $(date -Iseconds) ===" >> "$LOG"
check_one() {
local agent=$1
local cdp_port=$2
# node == agent == profile (unified naming)
check_agent "$agent" "$agent" "$cdp_port"
local rc=$?
if [ $rc -eq 0 ]; then
# Healthy: reset the consecutive-API-failure counter and close
# any open restart circuit.
rm -f "$STATE_DIR/failcount-$agent" "$STATE_DIR/futile-$agent" "$STATE_DIR/circuit-$agent" 2>/dev/null
return 0
fi
if [ $rc -eq 1 ]; then
# API failed but CDP port is listening: this is the false-kill vector
# (2026-10-04: single 30s API timeout killed 646's healthy browser).
# Require 2 CONSECUTIVE API failures before killing.
local count=0
local cf="$STATE_DIR/failcount-$agent"
[ -f "$cf" ] && count=$(cat "$cf" 2>/dev/null || echo 0)
count=$(( count + 1 ))
if [ "$count" -lt 2 ]; then
echo "$count" > "$cf"
echo "$(date -Iseconds) $agent: API failure $count of 2, deferring kill" >> "$LOG"
return 0
fi
rm -f "$cf" 2>/dev/null
else
# CDP port not listening (rc=2): zombie browser, kill immediately as before.
rm -f "$STATE_DIR/failcount-$agent" 2>/dev/null
fi
# De-conflict with chromebox-watchdog.sh: if the main browser for this
# profile launched <2 min ago, the watchdog just relaunched it — skip the
# kill path rather than racing it on a fresh cold start.
if recent_relaunch "$cdp_port"; then
echo "$(date -Iseconds) $agent: skipping kill (browser launched <2m ago, likely watchdog relaunch)" >> "$LOG"
return 0
fi
# Circuit breaker: repeated futile restarts stop here until cooldown.
if ! circuit_allows "$agent"; then
return 0
fi
restart_browser "$agent" "$cdp_port"
# 2026-10-04: post-restart re-check grace extended to ~60s total
# (restart_browser sleeps 15s internally + 45s here), matching
# chromebox-watchdog.sh's proven 60s retry window. Cold starts on a
# loaded box (load ~7) need more than 20s before CDP/API respond.
sleep 45
if ! check_agent "$agent" "$agent" "$cdp_port"; then
echo "$(date -Iseconds) $agent: CRITICAL - still down after restart" >> "$LOG"
circuit_note_restart "$agent" fail
else
echo "$(date -Iseconds) $agent: RECOVERED after restart" >> "$LOG"
circuit_note_restart "$agent" ok
fi
}
# Every active node in the NODES.md registry gets checked — new nodes
# propagate automatically, no per-node blocks to add.
"$NETVM_BIN/netvm-registry.py" 2>/dev/null | while IFS=: read -r node port; do
[ -n "$node" ] && [ -n "$port" ] && check_one "$node" "$port"
done
# Trim log (keep last 1000 lines)
tail -1000 "$LOG" > "$LOG.tmp" && mv "$LOG.tmp" "$LOG"
-905
View File
@@ -1,905 +0,0 @@
#!/usr/bin/env python3
"""agent-manager.py — Agent runs across tailnet devices (stdlib curses).
Read-only dashboard that aggregates agent harness runs (any CLI / bin
runtime: muse, agy, and whatever else matches the open taxonomy) over
SSH to every reachable tailnet device, then sorts by agent type.
Surfaces 3 read-only views (no actions in v1):
[1] RUNS — every run, sorted by agent type, then device.
[2] TYPES — counts per agent type with per-device breakdown.
[3] DEVICES — per-device reachability, run counts, and notes.
Structure mirrors bin/box-fleet-tui.py: all data-gathering lives in
pure, testable functions taking injected runners (see gather_*); the
curses UI is a thin renderer over those functions. The backend differs:
instead of local repo files, each refresh fans out over tailnet SSH
(one call per device: tmux panes + process table, joined locally).
Unreachable devices yield "n/a" rows — never a crash.
Usage:
python3 bin/agent-manager.py [--once [--json]] # non-interactive dump
"""
from __future__ import annotations
import concurrent.futures
import curses
import json
import os
import re
import subprocess
import sys
import time
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Callable, Dict, Iterable, List, Optional, Tuple
REPO_ROOT = Path(__file__).resolve().parent.parent
NA = "n/a"
RunFn = Callable[..., Tuple[int, str]]
# Per-device SSH budget; whole-fleet refresh runs devices in parallel.
DEVICE_TIMEOUT_S = 15
SSH_OPTS = ["-o", "BatchMode=yes", "-o", "ConnectTimeout=5"]
# Only these OS classes get an SSH probe (from `tailscale status`).
SSH_OS = ("linux", "macos")
# =====================================================================
# Agent-type taxonomy (open: unknown bin runtimes still show up)
# =====================================================================
#
# Classification keys off the argv[0] basename so a wrapper path never
# matters (/home/super/.local/bin/agy.bin -> agy). Scripts (*.py,
# *.sh) are never harness runtimes. Anything shaped like a bin runtime
# (*.bin, *-bin-*) that is not otherwise known shows under its own
# "other:<basename>" type instead of being dropped.
HARNESS_EXACT = {
"agy": "agy",
"agy.bin": "agy",
"muse-code": "muse",
"Muse": "muse",
"claude": "claude",
"codex": "codex",
"gemini": "gemini",
"aider": "aider",
"opencode": "opencode",
"crush": "crush",
"amp": "amp",
}
HARNESS_PREFIX = (
("muse-bin", "muse"),
)
# Bin-shaped names that are infrastructure, not agent harnesses.
HARNESS_DENY = frozenset({"tmux.bin", "tmux", "ssh.bin"})
SCRIPT_SUFFIXES = (".py", ".pyc", ".sh", ".pl", ".rb", ".js")
def classify_harness(argv0: str) -> Optional[Tuple[str, str]]:
"""Map an argv[0] to (agent_type, bin_name); None when not a harness.
Known harnesses collapse to a canonical type ("agy"); unknown bin
runtimes keep their own "other:<basename>" type so new harnesses
appear without a code change.
"""
base = os.path.basename((argv0 or "").strip().strip("'\""))
if not base:
return None
low = base.lower()
if low in HARNESS_DENY:
return None
if low.endswith(SCRIPT_SUFFIXES):
return None
if low in HARNESS_EXACT:
return HARNESS_EXACT[low], base
for prefix, typ in HARNESS_PREFIX:
if low.startswith(prefix):
return typ, base
if low.endswith(".bin") or "-bin-" in low or low.startswith("bin-"):
return "other:%s" % base, base
return None
def type_sort_key(agent_type: str) -> Tuple[int, str]:
"""Known types first (alpha), then other:* (alpha)."""
if agent_type.startswith("other:"):
return 1, agent_type
return 0, agent_type
# =====================================================================
# Default IO primitives (injectable seams for tests)
# =====================================================================
def _run(cmd: List[str], timeout: int = 15) -> Tuple[int, str]:
"""Run cmd, capture output. Returns (returncode, combined_output)."""
try:
r = subprocess.run(cmd, capture_output=True, text=True,
timeout=timeout)
return r.returncode, ((r.stdout or "") + (r.stderr or "")).strip()
except subprocess.TimeoutExpired:
return 124, "timed out after %ds: %s" % (timeout, " ".join(cmd))
except OSError as e:
return 127, str(e)
def _run_ssh(device: str, remote_cmd: str,
timeout: int = DEVICE_TIMEOUT_S,
run: Optional[RunFn] = None) -> Tuple[int, str]:
"""Run one remote command over tailnet SSH. Fails closed, never raises."""
run = run or _run
try:
return run(["ssh"] + SSH_OPTS + [device, remote_cmd],
timeout=timeout)
except Exception as e:
return 127, str(e)
# =====================================================================
# Pure parsers
# =====================================================================
def parse_tailscale_status(out: str) -> List[Dict[str, Any]]:
"""Parse `tailscale status` into [{name, ip, os, online, detail}].
Unparseable lines are skipped; a warning preamble is ignored.
"""
devices: List[Dict[str, Any]] = []
for line in (out or "").splitlines():
line = line.rstrip()
if not line or line.startswith("Warning:"):
continue
parts = line.split()
if len(parts) < 4:
continue
ip, name, _user, osname = parts[0], parts[1], parts[2], parts[3]
if not re.match(r"^\d+\.\d+\.\d+\.\d+$", ip):
continue
detail = " ".join(parts[4:]) if len(parts) > 4 else ""
online = not detail.startswith("offline")
devices.append({"name": name, "ip": ip, "os": osname.lower(),
"online": online, "detail": detail or NA})
return devices
PANE_PREFIX = "PANE|"
PS_MARKER = "__PS__"
def parse_panes(out: str) -> List[Dict[str, Any]]:
"""Parse tmux pane lines (PANE|sock|id|pid|session|cmd|title).
Pane ids repeat across servers, so (socket, id) is the identity.
Session groups repeat a pane under several sessions; dedupe by
(socket, id), joining session names. The legacy 4-field shape
(no socket) still parses with sock="".
"""
seen: Dict[Tuple[str, str], Dict[str, Any]] = {}
order: List[Tuple[str, str]] = []
for line in (out or "").splitlines():
if not line.startswith(PANE_PREFIX):
continue
fields = line[len(PANE_PREFIX):].split("|")
if len(fields) >= 6:
sock, pane_id, pid_s, session, cmd = (
fields[0].strip(), fields[1].strip(), fields[2].strip(),
fields[3].strip(), fields[4].strip())
title = "|".join(fields[5:]).strip()
elif len(fields) >= 4:
sock, pane_id, pid_s, session, cmd = (
"", fields[0].strip(), fields[1].strip(),
fields[2].strip(), fields[3].strip())
title = "|".join(fields[4:]).strip()
else:
continue
try:
pid = int(pid_s)
except ValueError:
continue
session = session or NA
key = (sock, pane_id)
if key in seen:
prev = seen[key]["session"]
if session not in prev.split(","):
seen[key]["session"] = prev + "," + session
continue
seen[key] = {"sock": sock, "pane": pane_id, "pid": pid,
"session": session, "cmd": cmd, "title": title}
order.append(key)
return [seen[k] for k in order]
def display_session(run: Dict[str, Any]) -> str:
"""Session label; socket-qualified unless it is the default server."""
session = str(run.get("session", NA))
sock = run.get("sock") or ""
if session == NA or sock in ("", "default"):
return session
return "%s:%s" % (sock, session)
def parse_ps(out: str) -> Dict[int, Dict[str, Any]]:
"""Parse `ps -eo pid,ppid,etime,command` into {pid: rec}."""
procs: Dict[int, Dict[str, Any]] = {}
for line in (out or "").splitlines():
parts = line.split(None, 3)
if len(parts) < 4:
continue
try:
pid, ppid = int(parts[0]), int(parts[1])
except ValueError:
continue # header row
procs[pid] = {"pid": pid, "ppid": ppid, "etime": parts[2],
"args": parts[3]}
return procs
def split_scan(out: str) -> Tuple[str, str]:
"""Split a device scan into (pane_text, ps_text) at the marker."""
if PS_MARKER in out:
pane_text, _, ps_text = out.partition(PS_MARKER)
return pane_text, ps_text
return out, ""
def _descendants(procs: Dict[int, Dict[str, Any]],
root: int) -> List[int]:
"""Pids under root (breadth-first via ppid links)."""
kids: Dict[int, List[int]] = {}
for pid, rec in procs.items():
kids.setdefault(rec["ppid"], []).append(pid)
out: List[int] = []
queue = list(kids.get(root, []))
seen = {root}
while queue:
pid = queue.pop(0)
if pid in seen:
continue
seen.add(pid)
out.append(pid)
queue.extend(kids.get(pid, []))
return out
def join_runs(panes: List[Dict[str, Any]],
procs: Dict[int, Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Join tmux panes with harness processes into run records.
A pane is a run when its current command classifies as a harness
or a harness binary runs among its descendants. Harness processes
under no pane surface as bare runs (session/pane n/a). Returns
records sorted by (agent_type, session, pane).
"""
runs: List[Dict[str, Any]] = []
claimed: set = set()
pane_roots = {p["pid"] for p in panes}
for pane in panes:
argv = (pane.get("cmd") or "").strip()
pane_hit = classify_harness(argv.split()[0] if argv else "")
# Always resolve the live harness descendant: the pane's root
# is usually a shell, so its pid/etime/bin would mislead.
# The pane-command match is only a fallback (stale command).
hit = None
hpid: Optional[int] = None
for pid in _descendants(procs, pane["pid"]):
rec = procs.get(pid)
if not rec:
continue
first = rec["args"].split()[0] if rec["args"] else ""
hit = classify_harness(first)
if hit is not None:
hpid = pid
break
if hit is None:
if pane_hit is None:
continue
hit, hpid = pane_hit, pane["pid"]
agent_type, _bin = hit
claimed.add(hpid)
prec = procs.get(hpid, {})
if prec.get("args"):
binn = os.path.basename(prec["args"].split()[0])
elif argv:
binn = argv.split()[0]
else:
binn = NA
runs.append({
"type": agent_type,
"bin": binn,
"sock": pane.get("sock", ""),
"session": pane.get("session", NA),
"pane": pane.get("pane", NA),
"pid": hpid,
"etime": prec.get("etime", NA),
"title": (pane.get("title") or "")[:48],
})
# Bare harness processes (no tmux pane above them).
for pid, rec in procs.items():
if pid in claimed:
continue
first = rec["args"].split()[0] if rec["args"] else ""
hit = classify_harness(first)
if hit is None:
continue
# Skip when some pane root is an ancestor (already covered).
anc, under_pane = rec["ppid"], False
hops = 0
while anc in procs and hops < 64:
if anc in pane_roots:
under_pane = True
break
anc = procs[anc]["ppid"]
hops += 1
if under_pane:
continue
claimed.add(pid)
runs.append({
"type": hit[0],
"bin": os.path.basename(first),
"sock": "",
"session": NA,
"pane": NA,
"pid": pid,
"etime": rec.get("etime", NA),
"title": "",
})
runs.sort(key=lambda r: (type_sort_key(r["type"]),
str(r["session"]), str(r["pane"])))
return runs
# =====================================================================
# Surfaces: devices + runs
# =====================================================================
PANE_FORMAT = ("PANE|%s|#{pane_id}|#{pane_pid}|#{session_name}|"
"#{pane_current_command}|#{pane_title}")
REMOTE_SCAN_CMD = (
"for d in \"${TMUX_TMPDIR:-/tmp}/tmux-$(id -u)\" "
"\"${TMPDIR:-/tmp}/tmux-$(id -u)\" /tmp/tmux-$(id -u); do "
"for s in \"$d\"/*; do [ -S \"$s\" ] || continue; "
"n=$(basename \"$s\"); "
"tmux -S \"$s\" list-panes -a -F \"PANE|$n|#{pane_id}|#{pane_pid}|"
"#{session_name}|#{pane_current_command}|#{pane_title}\" "
"2>/dev/null; done; done; "
"echo '%s'; ps -eo pid,ppid,etime,command 2>/dev/null" % PS_MARKER
)
def local_tmux_sockets() -> List[str]:
"""Absolute tmux socket paths on this machine (may be empty)."""
uid = os.getuid() if hasattr(os, "getuid") else 0
cands = []
for base in (os.environ.get("TMUX_TMPDIR") or "/tmp",
os.environ.get("TMPDIR") or "/tmp", "/tmp"):
cands.append(os.path.join(base, "tmux-%d" % uid))
found: List[str] = []
seen_dirs = set()
for d in cands:
if d in seen_dirs:
continue
seen_dirs.add(d)
try:
names = sorted(os.listdir(d))
except Exception:
continue
for n in names:
p = os.path.join(d, n)
try:
import stat
if stat.S_ISSOCK(os.stat(p).st_mode):
found.append(p)
except Exception:
continue
# Same server via two spellings: keep first per basename.
dedup: List[str] = []
seen_base = set()
for p in found:
b = os.path.basename(p)
if b not in seen_base:
seen_base.add(b)
dedup.append(p)
return dedup
def local_device_name() -> str:
"""Short hostname of this machine (never raises)."""
try:
import socket
return socket.gethostname().split(".")[0]
except Exception:
return "localhost"
def gather_devices(run: Optional[RunFn] = None,
local_name: Optional[str] = None) -> Dict[str, Any]:
"""Tailnet devices from `tailscale status` + the local machine.
Returns {"devices": [{name, ip, os, online, local, ssh, detail}],
"note": str}. Devices are stable-sorted: local first, then by name.
"ssh" marks whether v1 probes the device (online + ssh-capable OS).
"""
run = run or _run
local_name = local_name or local_device_name()
rc, out = run(["tailscale", "status"], timeout=10)
if rc != 0:
return {"devices": [{"name": local_name, "ip": NA, "os": NA,
"online": True, "local": True, "ssh": False,
"detail": "local only"}],
"note": "tailscale status failed (%s); local only."
% (out.strip().splitlines()[-1][:80] if out.strip()
else "rc=%d" % rc)}
devices = []
for d in parse_tailscale_status(out):
local = (d["name"] == local_name)
ssh = bool(d["online"] and d["os"] in SSH_OS and not local)
devices.append({"name": d["name"], "ip": d["ip"], "os": d["os"],
"online": d["online"], "local": local, "ssh": ssh,
"detail": d["detail"]})
if not any(d["local"] for d in devices):
devices.append({"name": local_name, "ip": NA, "os": NA,
"online": True, "local": True, "ssh": False,
"detail": "local"})
devices.sort(key=lambda d: (not d["local"], d["name"]))
return {"devices": devices, "note": ""}
def gather_device_runs(device: Dict[str, Any],
run: Optional[RunFn] = None) -> Dict[str, Any]:
"""One device scan -> {device, ok, runs, note}. Never raises."""
run = run or _run
name = device.get("name", "?")
try:
if device.get("local"):
pane_chunks = []
for sock in local_tmux_sockets():
rc1, chunk = run(
["tmux", "-S", sock, "list-panes", "-a", "-F",
PANE_FORMAT % os.path.basename(sock)],
timeout=DEVICE_TIMEOUT_S)
if rc1 == 0 and chunk:
pane_chunks.append(chunk)
_rc2, ps_out = run(["ps", "-eo", "pid,ppid,etime,command"],
timeout=DEVICE_TIMEOUT_S)
out = ("\n".join(pane_chunks) + "\n" + PS_MARKER + "\n"
+ ps_out)
else:
rc, out = _run_ssh(name, REMOTE_SCAN_CMD, run=run)
ssh_err = "" if rc == 0 else out.strip().splitlines()
ssh_err = ssh_err[-1][:100] if ssh_err else "rc=%d" % rc
if rc != 0:
return {"device": name, "ok": False, "runs": [],
"note": "ssh failed: %s" % ssh_err}
pane_text, ps_text = split_scan(out)
runs = join_runs(parse_panes(pane_text), parse_ps(ps_text))
for r in runs:
r["device"] = name
return {"device": name, "ok": True, "runs": runs, "note": ""}
except Exception as e:
return {"device": name, "ok": False, "runs": [],
"note": "scan failed: %s" % e}
def gather_all(run: Optional[RunFn] = None,
devices: Optional[List[Dict[str, Any]]] = None,
max_workers: int = 8) -> Dict[str, Any]:
"""One-shot snapshot: devices + runs sorted by agent type.
Devices scan in parallel (threads); each device is isolated — one
failure never blocks the rest. Returns {"devices": [...],
"runs": [...] (sorted by type/device), "by_type": {type: {total,
devices: {name: n}}}, "unreachable": [names], "note": str}.
"""
run = run or _run
dev_info = gather_devices(run=run)
if devices is None:
devices = [d for d in dev_info["devices"]
if d.get("local") or d.get("ssh")]
else:
devices = [d for d in devices
if d.get("local") or d.get("ssh")]
probed = {d["name"] for d in devices}
results: List[Dict[str, Any]] = []
if devices:
with concurrent.futures.ThreadPoolExecutor(
max_workers=min(max_workers, len(devices))) as pool:
futs = {pool.submit(gather_device_runs, d, run): d["name"]
for d in devices}
for fut in concurrent.futures.as_completed(futs):
try:
results.append(fut.result())
except Exception as e:
results.append({"device": futs[fut], "ok": False,
"runs": [], "note": "scan error: %s" % e})
runs: List[Dict[str, Any]] = []
unreachable: List[str] = []
for res in results:
if not res.get("ok"):
unreachable.append(res["device"])
continue
runs.extend(res.get("runs", []))
runs.sort(key=lambda r: (type_sort_key(r["type"]), r.get("device", ""),
str(r.get("session", ""))))
by_type: Dict[str, Dict[str, Any]] = {}
for r in runs:
bucket = by_type.setdefault(r["type"], {"total": 0, "devices": {}})
bucket["total"] += 1
dev = r.get("device", "?")
bucket["devices"][dev] = bucket["devices"].get(dev, 0) + 1
skipped = sorted(d["name"] for d in dev_info["devices"]
if d["name"] not in probed)
notes = [dev_info["note"]] if dev_info["note"] else []
if skipped:
notes.append("skipped (offline/mobile/key): %s" % ", ".join(skipped))
return {"devices": dev_info["devices"], "runs": runs,
"by_type": by_type, "unreachable": sorted(unreachable),
"note": "; ".join(notes)}
# =====================================================================
# Curses UI (thin read-only renderer over gather_*)
# =====================================================================
AUTO_REFRESH_S = 30.0
class AgentManagerTUI:
"""Read-only agent-run console. q quits, r refreshes, ? helps."""
def __init__(self, stdscr: "curses.window"):
self.stdscr = stdscr
self.current_tab = 0
self.tabs = [
"1: RUNS",
"2: TYPES",
"3: DEVICES",
]
try:
curses.curs_set(0)
except Exception:
pass
self.stdscr.nodelay(True)
self.stdscr.keypad(True)
if hasattr(curses, "set_escdelay"):
try:
curses.set_escdelay(25)
except Exception:
pass
self._init_colors()
self.scroll = 0
self.show_help = False
self.status_msg = "Scanning tailnet devices..."
self.last_refresh = 0.0
self.snapshot: Dict[str, Any] = {}
self.refresh()
# -- setup ------------------------------------------------------
def _init_colors(self) -> None:
try:
curses.start_color()
curses.use_default_colors()
curses.init_pair(1, curses.COLOR_CYAN, -1)
curses.init_pair(2, curses.COLOR_YELLOW, -1)
curses.init_pair(3, curses.COLOR_GREEN, -1)
curses.init_pair(4, curses.COLOR_RED, -1)
curses.init_pair(5, curses.COLOR_MAGENTA, -1)
curses.init_pair(6, curses.COLOR_BLACK, curses.COLOR_CYAN)
curses.init_pair(7, curses.COLOR_BLACK, curses.COLOR_WHITE)
curses.init_pair(8, curses.COLOR_BLACK, curses.COLOR_YELLOW)
except Exception:
pass
def _attr(self, name: str) -> int:
try:
mapping = {
"normal": curses.color_pair(0),
"cyan": curses.color_pair(1) | curses.A_BOLD,
"yellow": curses.color_pair(2) | curses.A_BOLD,
"green": curses.color_pair(3) | curses.A_BOLD,
"red": curses.color_pair(4) | curses.A_BOLD,
"magenta": curses.color_pair(5) | curses.A_BOLD,
"head_sel": curses.color_pair(6) | curses.A_BOLD,
"row_sel": curses.color_pair(7) | curses.A_BOLD,
"warn": curses.color_pair(8) | curses.A_BOLD,
"dim": curses.A_DIM,
}
return mapping.get(name, 0)
except Exception:
return 0
# -- data -------------------------------------------------------
def refresh(self) -> None:
try:
self.snapshot = gather_all()
runs = len(self.snapshot.get("runs", []))
devs = len([d for d in self.snapshot.get("devices", [])
if d.get("local") or d.get("ssh")])
self.status_msg = (
"Snapshot %s: %d runs on %d devices "
"(auto-refresh %ds; r=refresh)" % (
datetime.now().strftime("%H:%M:%S"), runs, devs,
int(AUTO_REFRESH_S)))
except Exception as e:
self.snapshot = {}
self.status_msg = "Refresh failed (showing n/a): %s" % e
self.last_refresh = time.time()
self.scroll = 0
# -- render helpers ---------------------------------------------
def safe_addstr(self, y: int, x: int, text: str, attr: int = 0) -> None:
h, w = self.stdscr.getmaxyx()
if 0 <= y < h and 0 <= x < w:
try:
self.stdscr.addstr(y, x, text[:max(0, w - x - 1)], attr)
except Exception:
pass
def _render_header(self, w: int) -> None:
self.safe_addstr(0, 0, " " * w, self._attr("head_sel"))
title = " AGENT MANAGER (read-only) [agent-manager.py] "
self.safe_addstr(0, 1, title, self._attr("head_sel"))
self.safe_addstr(1, 0, " " * w, self._attr("dim"))
col = 1
for idx, tab_name in enumerate(self.tabs):
pill = " [%s] " % tab_name
attr = self._attr("head_sel") if idx == self.current_tab \
else self._attr("dim")
self.safe_addstr(1, col, pill, attr)
col += len(pill) + 1
self.safe_addstr(2, 0, "-" * w, self._attr("dim"))
def _render_footer(self, h: int, w: int) -> None:
self.safe_addstr(h - 2, 0, "-" * w, self._attr("dim"))
hints = " 1-3/Tab: Tabs j/k: Scroll r: Refresh ?: Help q: Quit"
self.safe_addstr(h - 1, 1, self.status_msg[: w - 2],
self._attr("dim"))
if len(self.status_msg) + len(hints) + 2 < w:
self.safe_addstr(h - 1, w - len(hints) - 1, hints,
self._attr("dim"))
def _body(self, h: int, w: int, title: str,
lines: List[Tuple[str, str]]) -> None:
self.safe_addstr(3, 2, title, self._attr("cyan"))
self.safe_addstr(4, 2, "-" * (w - 4), self._attr("dim"))
max_rows = h - 8
visible = lines[self.scroll:self.scroll + max_rows]
for i, (text, attr_name) in enumerate(visible):
self.safe_addstr(5 + i, 2, text, self._attr(attr_name))
if self.scroll > 0:
self.safe_addstr(5, w - 6, "^more", self._attr("dim"))
if self.scroll + max_rows < len(lines):
self.safe_addstr(h - 3, w - 6, "vmore", self._attr("dim"))
def _note_lines(self, w: int) -> List[Tuple[str, str]]:
note = self.snapshot.get("note", "")
unreach = self.snapshot.get("unreachable", [])
lines: List[Tuple[str, str]] = []
if unreach:
lines.append(("", "normal"))
lines.append(("unreachable: %s" % ", ".join(unreach),
"red"))
if note:
lines.append(("", "normal"))
lines.append(("note: %s" % note[: w - 10], "yellow"))
return lines
# -- per-tab renderers ------------------------------------------
def _render_runs(self, h: int, w: int) -> None:
runs = self.snapshot.get("runs", [])
lines: List[Tuple[str, str]] = [
("%-14s %-12s %-16s %-14s %-6s %-11s %s"
% ("TYPE", "DEVICE", "BIN", "SESSION", "PANE", "ELAPSED",
"TITLE"), "dim"),
]
last_type = None
for r in runs:
typ = r.get("type", "?")
if typ != last_type:
lines.append(("", "normal"))
last_type = typ
attr = "green" if not typ.startswith("other:") else "yellow"
lines.append((
"%-14s %-12s %-16s %-14s %-6s %-11s %s" % (
typ[:14], r.get("device", "?")[:12],
r.get("bin", NA)[:16], display_session(r)[:14],
str(r.get("pane", NA))[:6],
str(r.get("etime", NA))[:11],
r.get("title", "")[: w - 80]), attr))
if not runs:
lines.append(("(No agent runs found on probed devices.)",
"dim"))
lines.extend(self._note_lines(w))
self._body(h, w, "AGENT RUNS SORTED BY TYPE (%d)" % len(runs),
lines)
def _render_types(self, h: int, w: int) -> None:
by_type = self.snapshot.get("by_type", {})
lines: List[Tuple[str, str]] = []
total = sum(b.get("total", 0) for b in by_type.values())
lines.append(("Agent types: %d | total runs: %d"
% (len(by_type), total), "cyan"))
lines.append(("", "normal"))
for typ in sorted(by_type, key=type_sort_key):
bucket = by_type[typ]
attr = "green" if not typ.startswith("other:") else "yellow"
lines.append(("%-16s %d" % (typ, bucket.get("total", 0)),
attr))
for dev, n in sorted(bucket.get("devices", {}).items()):
lines.append((" %-14s %d" % (dev, n), "normal"))
lines.append(("", "normal"))
if not by_type:
lines.append(("(No agent types observed.)", "dim"))
lines.extend(self._note_lines(w))
self._body(h, w, "COUNTS BY AGENT TYPE", lines)
def _render_devices(self, h: int, w: int) -> None:
devices = self.snapshot.get("devices", [])
runs = self.snapshot.get("runs", [])
counts: Dict[str, int] = {}
for r in runs:
dev = r.get("device", "?")
counts[dev] = counts.get(dev, 0) + 1
unreach = set(self.snapshot.get("unreachable", []))
lines: List[Tuple[str, str]] = [
("%-24s %-15s %-7s %-7s %-5s %s"
% ("DEVICE", "IP", "OS", "PROBED", "RUNS", "DETAIL"), "dim"),
]
for d in devices:
name = d.get("name", "?")
probed = bool(d.get("local") or d.get("ssh"))
if name in unreach:
attr = "red"
elif not probed:
attr = "dim"
elif counts.get(name):
attr = "green"
else:
attr = "normal"
lines.append((
"%-24s %-15s %-7s %-7s %-5s %s" % (
name[:24] + (" *" if d.get("local") else ""),
d.get("ip", NA)[:15], str(d.get("os", NA))[:7],
"yes" if probed else "no",
counts.get(name, 0) if probed else NA,
str(d.get("detail", ""))[: w - 66]), attr))
lines.append(("", "normal"))
lines.append(("* = local (no SSH); unreachable shows n/a, never "
"blocks the rest.", "dim"))
lines.extend(self._note_lines(w))
self._body(h, w, "TAILNET DEVICES", lines)
def _render_help(self, h: int, w: int) -> None:
modal_w = min(64, w - 6)
modal_h = 13
top = (h - modal_h) // 2
left = (w - modal_w) // 2
for y in range(top, top + modal_h):
self.safe_addstr(y, left, " " * modal_w, self._attr("row_sel"))
self.safe_addstr(top, left, "+" + "-" * (modal_w - 2) + "+",
self._attr("cyan"))
for y in range(top + 1, top + modal_h - 1):
self.safe_addstr(y, left, "|", self._attr("cyan"))
self.safe_addstr(y, left + modal_w - 1, "|",
self._attr("cyan"))
self.safe_addstr(top + modal_h - 1, left,
"+" + "-" * (modal_w - 2) + "+",
self._attr("cyan"))
self.safe_addstr(top + 1, left + 3, "AGENT MANAGER HELP (read-only)",
self._attr("cyan"))
for i, hint in enumerate([
"1-3 / Tab: switch surfaces",
"j/k / Up/Down: scroll",
"r: refresh snapshot now",
"q / Esc: quit (Esc closes help first)",
"",
"One SSH call per device per refresh.",
"Missing data shows as 'n/a' — never a crash.",
]):
self.safe_addstr(top + 3 + i, left + 4, hint,
self._attr("normal"))
# -- input + main loop ------------------------------------------
def _handle_key(self, ch: int) -> bool:
if ch in (3, 4): # Ctrl+C / Ctrl+D
return False
if self.show_help:
if ch in (27, ord("q"), ord("Q"), ord("?")):
self.show_help = False
return True
if ch in (ord("q"), ord("Q")):
return False
if ch in (27, ord("?")):
self.show_help = True
return True
if ch in (ord("1"), ord("2"), ord("3")):
self.current_tab = ch - ord("1")
self.scroll = 0
return True
if ch == ord("\t"):
self.current_tab = (self.current_tab + 1) % len(self.tabs)
self.scroll = 0
return True
if ch in (ord("j"), curses.KEY_DOWN):
self.scroll += 1
return True
if ch in (ord("k"), curses.KEY_UP):
self.scroll = max(0, self.scroll - 1)
return True
if ch in (ord("r"), ord("R")):
self.refresh()
return True
return True
def run(self) -> None:
renderers = [self._render_runs, self._render_types,
self._render_devices]
while True:
h, w = self.stdscr.getmaxyx()
self.stdscr.erase()
self._render_header(w)
try:
renderers[self.current_tab](h, w)
except Exception as e:
self.safe_addstr(5, 4, "Render error (n/a): %s" % e,
self._attr("red"))
self._render_footer(h, w)
if self.show_help:
self._render_help(h, w)
self.stdscr.refresh()
try:
ch = self.stdscr.getch()
if ch != -1 and not self._handle_key(ch):
break
except KeyboardInterrupt:
break
if time.time() - self.last_refresh > AUTO_REFRESH_S:
self.refresh()
time.sleep(0.05)
def main(argv: Optional[List[str]] = None) -> int:
argv = list(sys.argv[1:] if argv is None else argv)
if "--once" in argv:
snap = gather_all()
if "--json" in argv:
print(json.dumps(snap, indent=2, default=str))
else:
print("== RUNS (%d) ==" % len(snap.get("runs", [])))
for r in snap.get("runs", []):
print("%-14s %-12s %-16s %s/%s %s" % (
r.get("type"), r.get("device"), r.get("bin"),
display_session(r), r.get("pane"), r.get("etime")))
print("== BY TYPE == ")
for typ, b in sorted(snap.get("by_type", {}).items()):
print("%s: %d %s" % (typ, b["total"], b["devices"]))
print("unreachable: %s" % snap.get("unreachable"))
if snap.get("note"):
print("note: %s" % snap["note"])
return 0
curses.wrapper(lambda stdscr: AgentManagerTUI(stdscr).run())
return 0
if __name__ == "__main__":
sys.exit(main())
+1
View File
@@ -0,0 +1 @@
agent-cognitive-probe.py
-625
View File
@@ -1,625 +0,0 @@
#!/usr/bin/env python3
"""agent_md.py — Access and modify Muse agent .md files via Hatch gateway and SSH.
Enables operators to inspect, audit, diff, and inject operational DRIVE into
Muse agents across the fleet (muse, pip, 646, opm, def, dev).
"""
import difflib
import json
import os
import re
import sys
import time
from pathlib import Path
# Ensure muse_cli can be imported
sys.path.insert(0, os.path.expanduser("~/.local/lib/python3.14/site-packages"))
try:
from muse_cli.gateway import Gateway, load_cookies, AuthError, GatewayError
except ImportError:
Gateway = None
NETVM_ROOT = Path("/home/super/Projects/NetVM")
SHARED_OPERATORS = NETVM_ROOT / "shared" / "operators"
VALID_ACCOUNTS = ["muse", "pip", "646", "opm", "def", "dev"]
TARGET_MD_FILES = [
"SOUL.md",
"PROACTIVE_PREFERENCES.md",
"HEARTBEAT.md",
"AGENTS.md",
"MEMORY.md",
"USER.md",
"TOOLS.md",
"IDENTITY.md",
]
class MDValidationError(ValueError):
"""An md account/filename/path failed safety validation.
box-ctl.py maps this to BAD_NAME; it is always raised before any
gateway call or filesystem write.
"""
MD_ACCOUNT_RE = re.compile(r"^[A-Za-z0-9][A-Za-z0-9_-]{0,31}$")
MD_FILENAME_RE = re.compile(r"^[A-Za-z0-9_.-]{1,128}$")
MD_SUBPATH_RE = re.compile(r"^[A-Za-z0-9_.-]+(/[A-Za-z0-9_.-]+)*$")
def validate_account(account: str) -> str:
"""Reject account values that could escape the cookies/config path."""
if not isinstance(account, str) or not MD_ACCOUNT_RE.fullmatch(account):
raise MDValidationError(
"Invalid agent account %r: must match ^[A-Za-z0-9][A-Za-z0-9_-]{0,31}$"
% (account,))
return account
def validate_filename(filename: str, template_only: bool = False) -> str:
"""Reject filenames that could escape the md directory.
Template flows (diff/amend/append/pull) additionally require one of
TARGET_MD_FILES, since they index into shared/operators/.
"""
if template_only:
if filename not in TARGET_MD_FILES:
raise MDValidationError(
"Unknown shared template %r: must be one of %s"
% (filename, sorted(TARGET_MD_FILES)))
return filename
if not isinstance(filename, str) or filename in (".", "..") \
or not MD_FILENAME_RE.fullmatch(filename):
raise MDValidationError(
"Invalid filename %r: plain basename, no directories" % (filename,))
return filename
def validate_subpath(path: str) -> str:
"""Reject list paths that escape the container root ('' = root)."""
if path in (None, ""):
return ""
if not isinstance(path, str) or not MD_SUBPATH_RE.fullmatch(path) \
or ".." in path.split("/"):
raise MDValidationError(
"Invalid list path %r: subdir without '..'" % (path,))
return path
# Tunnel / Port inventory
TUNNEL_PORTS = {
"muse-main": {"port": 2224, "terminal": 7681, "user": "muse"},
"muse": {"port": 2225, "terminal": 7682, "user": "hatch"},
"646": {"port": 2226, "terminal": 7683, "user": "hatch"},
"pip": {"port": 2227, "terminal": 7684, "user": "hatch"},
"opm": {"port": 2228, "terminal": 7685, "user": "hatch"},
"def": {"port": 2229, "terminal": 7686, "user": "hatch"},
"dev": {"port": 2230, "terminal": 7687, "user": "hatch"},
}
def get_gateway(account: str) -> "Gateway":
"""Obtain an authenticated Gateway connection for an account."""
validate_account(account)
if not Gateway:
raise RuntimeError("muse_cli.gateway module is not available")
conf_dir = Path.home() / ".config" / "muse-cli" / account
cfile = conf_dir / "cookies.txt"
if not cfile.exists():
raise FileNotFoundError(f"No cookies found for account '{account}' at {cfile}")
cookies = load_cookies(str(cfile))
if not cookies.strip():
raise ValueError(f"Cookies file for '{account}' is empty")
return Gateway(cookies)
def list_files(account: str, path: str = "") -> list:
"""List files in the agent container filesystem via Hatch."""
path = validate_subpath(path)
gw = get_gateway(account)
res = gw.call_json("fs.list", body={"path": path})
return res.get("entries", [])
def read_md(account: str, filename: str, max_bytes: int = 200000) -> dict:
"""Read a markdown file from the agent container via Hatch."""
validate_filename(filename)
gw = get_gateway(account)
offset = 0
chunks = []
chunk_len = min(65536, max_bytes)
while True:
body = {"path": filename, "offset": offset, "len": chunk_len}
res = gw.call_json("fs.read", body=body)
text = res.get("text", "")
if not text and res.get("data_base64"):
import base64
text = base64.b64decode(res["data_base64"]).decode("utf-8", "replace")
chunks.append(text)
offset += len(text.encode("utf-8"))
if res.get("eof") or offset >= max_bytes or not text:
break
full_text = "".join(chunks)
return {
"ok": True,
"account": account,
"filename": filename,
"size": len(full_text.encode("utf-8")),
"text": full_text,
"eof": True,
}
def write_md(account: str, filename: str, text: str, overwrite: bool = True, append: bool = False) -> dict:
"""Write content to a file in the agent container via Hatch."""
validate_filename(filename)
gw = get_gateway(account)
body = {
"path": filename,
"overwrite": overwrite,
"append": append,
"create_parent": True,
"text": text,
}
res = gw.call_json("fs.write", body=body)
return {
"ok": True,
"account": account,
"filename": filename,
"bytes_written": res.get("bytes_written", len(text.encode("utf-8"))),
"path": res.get("path", f"/{filename}"),
}
def audit_agents(accounts: list = None) -> dict:
"""Audit markdown files and operational DRIVE across fleet agents."""
accounts = accounts or VALID_ACCOUNTS
for acct in accounts:
validate_account(acct)
results = {}
for acct in accounts:
acct_res = {
"account": acct,
"connected": False,
"files": {},
"drive_status": {},
"issues": [],
"drive_score": 0,
}
try:
gw = get_gateway(acct)
acct_res["connected"] = True
acct_res["vm_id"] = gw.vm_id
entries = gw.call_json("fs.list", body={"path": ""}).get("entries", [])
entry_map = {e["name"]: e for e in entries}
for tf in TARGET_MD_FILES:
if tf in entry_map:
info = entry_map[tf]
acct_res["files"][tf] = {
"exists": True,
"size": info.get("size", 0),
"modified": info.get("modifiedAt", ""),
}
else:
acct_res["files"][tf] = {
"exists": False,
"size": 0,
"modified": None,
}
# Analyze DRIVE indicators
# 1. HEARTBEAT.md checklist
hb_info = acct_res["files"].get("HEARTBEAT.md", {})
if not hb_info.get("exists"):
acct_res["drive_status"]["heartbeat"] = "MISSING"
acct_res["issues"].append("HEARTBEAT.md missing (no recurring checks)")
elif hb_info.get("size", 0) <= 120:
acct_res["drive_status"]["heartbeat"] = "EMPTY_CHECKLIST"
acct_res["issues"].append("HEARTBEAT.md has empty checklist (background runner idle)")
else:
acct_res["drive_status"]["heartbeat"] = "ACTIVE"
acct_res["drive_score"] += 25
# 2. PROACTIVE_PREFERENCES.md
pro_info = acct_res["files"].get("PROACTIVE_PREFERENCES.md", {})
if not pro_info.get("exists"):
acct_res["drive_status"]["proactive"] = "MISSING"
acct_res["issues"].append("PROACTIVE_PREFERENCES.md missing")
elif pro_info.get("size", 0) <= 500:
acct_res["drive_status"]["proactive"] = "BLANK_TEMPLATE"
acct_res["issues"].append("PROACTIVE_PREFERENCES.md unconfigured (never reaches out)")
else:
acct_res["drive_status"]["proactive"] = "CONFIGURED"
acct_res["drive_score"] += 25
# 3. SOUL.md
soul_info = acct_res["files"].get("SOUL.md", {})
if not soul_info.get("exists"):
acct_res["drive_status"]["soul"] = "MISSING"
acct_res["issues"].append("SOUL.md missing")
elif soul_info.get("size", 0) <= 850:
acct_res["drive_status"]["soul"] = "PASSIVE_STOCK"
acct_res["issues"].append("SOUL.md is passive stock template (no operator drive)")
else:
acct_res["drive_status"]["soul"] = "OPERATOR_SOUL"
acct_res["drive_score"] += 25
# 4. TOOLS.md & USER.md
tools_info = acct_res["files"].get("TOOLS.md", {})
user_info = acct_res["files"].get("USER.md", {})
if tools_info.get("size", 0) > 400 and user_info.get("size", 0) > 400:
acct_res["drive_status"]["context"] = "FULL_CONTEXT"
acct_res["drive_score"] += 25
else:
acct_res["drive_status"]["context"] = "PARTIAL_OR_EMPTY"
acct_res["issues"].append("TOOLS.md or USER.md missing operational conventions")
except Exception as e:
acct_res["error"] = str(e)
acct_res["issues"].append(f"Connection failed: {e}")
results[acct] = acct_res
return results
def diff_md(account: str, filename: str) -> dict:
"""Compare an agent's container file against the shared operator template."""
validate_filename(filename, template_only=True)
local_path = SHARED_OPERATORS / filename
if not local_path.exists():
raise FileNotFoundError(f"Local template {local_path} not found")
local_content = local_path.read_text(encoding="utf-8")
remote_data = read_md(account, filename)
remote_content = remote_data.get("text", "")
diff = list(difflib.unified_diff(
remote_content.splitlines(keepends=True),
local_content.splitlines(keepends=True),
fromfile=f"{account}:{filename} (remote)",
tofile=f"shared/operators/{filename} (local)",
))
return {
"ok": True,
"account": account,
"filename": filename,
"identical": len(diff) == 0,
"diff": "".join(diff),
"remote_size": len(remote_content.encode("utf-8")),
"local_size": len(local_content.encode("utf-8")),
}
def amend_md(filename: str, content: str, author: str = "operator", reason: str = "") -> dict:
"""Amend a centralized shared operator template in shared/operators/ with safety validation and git commit."""
import subprocess
validate_filename(filename, template_only=True)
local_path = SHARED_OPERATORS / filename
if not local_path.exists():
raise FileNotFoundError(f"Shared operator file {filename} does not exist in {SHARED_OPERATORS}")
# Drive safety validation
if filename == "HEARTBEAT.md":
# Ensure checklist is not gutted
non_comment_lines = [l for l in content.splitlines() if l.strip() and not l.strip().startswith("#")]
checklist_items = [l for l in non_comment_lines if l.strip().startswith("- [")]
if not checklist_items:
raise ValueError("Safety rejection: amendment removes all active checklist items from HEARTBEAT.md")
elif filename == "PROACTIVE_PREFERENCES.md":
if len(content.strip()) < 400:
raise ValueError("Safety rejection: amendment would reduce PROACTIVE_PREFERENCES.md to unconfigured state")
elif filename == "SOUL.md":
if "Be a guest in someone's life" in content and "AUTONOMOUS OPERATIONAL DRIVE" not in content:
raise ValueError("Safety rejection: amendment reverts SOUL.md to passive stock template")
old_content = local_path.read_text(encoding="utf-8")
local_path.write_text(content, encoding="utf-8")
# Git auto-commit if in git repo
git_committed = False
git_hash = None
try:
commit_msg = f"amend(operators): update {filename} via {author}"
if reason:
commit_msg += f" - {reason}"
subprocess.run(["git", "add", str(local_path)], cwd=str(NETVM_ROOT), check=True, capture_output=True)
cr = subprocess.run(["git", "commit", "-m", commit_msg], cwd=str(NETVM_ROOT), capture_output=True, text=True)
if cr.returncode == 0:
git_committed = True
hr = subprocess.run(["git", "rev-parse", "--short", "HEAD"], cwd=str(NETVM_ROOT), capture_output=True, text=True)
git_hash = hr.stdout.strip()
except Exception:
pass
return {
"ok": True,
"action": "amend",
"filename": filename,
"author": author,
"reason": reason,
"bytes_written": len(content.encode("utf-8")),
"git_committed": git_committed,
"commit": git_hash,
}
def append_md(filename: str, text: str, author: str = "operator", section: str = None) -> dict:
"""Safely append an amendment or lesson to a centralized shared template."""
validate_filename(filename, template_only=True)
local_path = SHARED_OPERATORS / filename
if not local_path.exists():
raise FileNotFoundError(f"Shared operator file {filename} does not exist in {SHARED_OPERATORS}")
current = local_path.read_text(encoding="utf-8")
header = f"\n\n<!-- Amendment by {author} on {time.strftime('%Y-%m-%d %H:%M:%S UTC', time.gmtime())} -->\n"
if section:
header += f"### {section}\n"
new_content = current.rstrip() + header + text.strip() + "\n"
return amend_md(filename, new_content, author=author, reason=f"append {section or 'note'}")
def pull_md(account: str, filename: str) -> dict:
"""Pull the canonical centralized template from shared/operators/ into an agent's container."""
validate_filename(filename, template_only=True)
local_path = SHARED_OPERATORS / filename
if not local_path.exists():
raise FileNotFoundError(f"Shared operator file {filename} does not exist in {SHARED_OPERATORS}")
content = local_path.read_text(encoding="utf-8")
res = write_md(account, filename, content, overwrite=True)
return {
"ok": True,
"account": account,
"filename": filename,
"bytes_written": res.get("bytes_written"),
"message": f"Successfully pulled canonical {filename} into {account} container",
}
def inject_drive(account: str, force: bool = False) -> dict:
"""Inject high-drive operational instructions into the agent's container."""
updates = []
# 1. SOUL.md
soul_text = (SHARED_OPERATORS / "SOUL.md").read_text(encoding="utf-8")
r_soul = write_md(account, "SOUL.md", soul_text, overwrite=True)
updates.append({"file": "SOUL.md", "bytes": r_soul["bytes_written"]})
# 2. PROACTIVE_PREFERENCES.md
pro_text = (SHARED_OPERATORS / "PROACTIVE_PREFERENCES.md").read_text(encoding="utf-8")
r_pro = write_md(account, "PROACTIVE_PREFERENCES.md", pro_text, overwrite=True)
updates.append({"file": "PROACTIVE_PREFERENCES.md", "bytes": r_pro["bytes_written"]})
# 3. HEARTBEAT.md
hb_text = (SHARED_OPERATORS / "HEARTBEAT.md").read_text(encoding="utf-8")
r_hb = write_md(account, "HEARTBEAT.md", hb_text, overwrite=True)
updates.append({"file": "HEARTBEAT.md", "bytes": r_hb["bytes_written"]})
# 4. USER.md
user_text = (SHARED_OPERATORS / "USER.md").read_text(encoding="utf-8")
r_user = write_md(account, "USER.md", user_text, overwrite=True)
updates.append({"file": "USER.md", "bytes": r_user["bytes_written"]})
# 5. TOOLS.md
tools_text = (SHARED_OPERATORS / "TOOLS.md").read_text(encoding="utf-8")
r_tools = write_md(account, "TOOLS.md", tools_text, overwrite=True)
updates.append({"file": "TOOLS.md", "bytes": r_tools["bytes_written"]})
# 6. AGENTS.md (preserve existing custom lessons if present)
agents_template = (SHARED_OPERATORS / "AGENTS.md").read_text(encoding="utf-8")
try:
remote_agents = read_md(account, "AGENTS.md").get("text", "")
if "## Lessons" in remote_agents and len(remote_agents) > len(agents_template):
# Extract custom lessons from remote and merge
custom_lessons = remote_agents.split("## Lessons", 1)[1]
merged_agents = agents_template.rstrip() + "\n\n## Lessons" + custom_lessons
r_agents = write_md(account, "AGENTS.md", merged_agents, overwrite=True)
else:
r_agents = write_md(account, "AGENTS.md", agents_template, overwrite=True)
except Exception:
r_agents = write_md(account, "AGENTS.md", agents_template, overwrite=True)
updates.append({"file": "AGENTS.md", "bytes": r_agents["bytes_written"]})
return {
"ok": True,
"account": account,
"action": "inject_drive",
"updates": updates,
"message": f"Successfully injected high-drive operator files into {account} container",
}
def get_ssh_info(account: str = None) -> dict:
"""Return SSH connection coordinates and reverse tunnel configuration."""
jump_host = "34.139.37.135"
if account:
entry = TUNNEL_PORTS.get(account, {"port": 2226, "terminal": 7683, "user": "hatch"})
port = entry["port"]
user = entry["user"]
cmd = f"ssh -o StrictHostKeyChecking=no -p {port} {user}@localhost"
proxy_cmd = f"ssh -o StrictHostKeyChecking=no -J super@{jump_host} -p {port} {user}@localhost"
return {
"account": account,
"jump_host": jump_host,
"port": port,
"container_user": user,
"terminal_port": entry.get("terminal"),
"direct_from_vm": cmd,
"jump_command": proxy_cmd,
"cat_example": f"cat file.md | {proxy_cmd} 'cat > /home/hatch/file.md'",
}
return {
"jump_host": jump_host,
"tunnels": TUNNEL_PORTS,
}
def main():
import argparse
parser = argparse.ArgumentParser(description="Manage Muse agent .md files via Hatch and SSH")
sub = parser.add_subparsers(dest="cmd")
p_audit = sub.add_parser("audit", help="Audit .md files and DRIVE across all agents")
p_audit.add_argument("accounts", nargs="*", help="Optional account filter")
p_audit.add_argument("--json", action="store_true", help="Output JSON")
p_list = sub.add_parser("list", help="List container files via Hatch")
p_list.add_argument("account", help="Agent account")
p_list.add_argument("path", nargs="?", default="", help="Subdirectory path")
p_read = sub.add_parser("read", help="Read a markdown file via Hatch")
p_read.add_argument("account", help="Agent account")
p_read.add_argument("filename", help="Filename (e.g. SOUL.md)")
p_write = sub.add_parser("write", help="Write a markdown file via Hatch")
p_write.add_argument("account", help="Agent account")
p_write.add_argument("filename", help="Filename (e.g. SOUL.md)")
p_write.add_argument("--content", help="Text content to write")
p_write.add_argument("--file", help="Local file to copy content from")
p_diff = sub.add_parser("diff", help="Diff remote file against shared operator template")
p_diff.add_argument("account", help="Agent account")
p_diff.add_argument("filename", help="Filename (e.g. SOUL.md)")
p_amend = sub.add_parser("amend", help="Amend a shared operator file with validation and git commit")
p_amend.add_argument("filename", help="Filename (e.g. AGENTS.md, TOOLS.md)")
p_amend.add_argument("--content", help="New content")
p_amend.add_argument("--file", help="File with new content")
p_amend.add_argument("--author", default="operator", help="Author of amendment")
p_amend.add_argument("--reason", default="", help="Reason for amendment")
p_append = sub.add_parser("append", help="Safely append a note or lesson to a shared operator file")
p_append.add_argument("filename", help="Filename (e.g. AGENTS.md)")
p_append.add_argument("text", help="Text to append")
p_append.add_argument("--author", default="operator", help="Author of amendment")
p_append.add_argument("--section", default=None, help="Optional section header")
p_pull = sub.add_parser("pull", help="Pull canonical shared file into an agent's container")
p_pull.add_argument("account", help="Agent account")
p_pull.add_argument("filename", help="Filename (e.g. SOUL.md)")
p_drive = sub.add_parser("inject-drive", help="Inject high-drive operator files into agent")
p_drive.add_argument("account", help="Agent account")
p_drive.add_argument("--force", action="store_true", help="Force overwrite")
p_sync_all = sub.add_parser("sync-all", help="Inject high-drive files across all active agents")
p_ssh = sub.add_parser("ssh-info", help="Get SSH tunnel dial-in information")
p_ssh.add_argument("account", nargs="?", help="Optional agent account")
args = parser.parse_args()
if not args.cmd:
parser.print_help()
sys.exit(1)
if args.cmd == "audit":
res = audit_agents(args.accounts or None)
if args.json:
print(json.dumps(res, indent=2))
else:
print(f"\n{'='*70}\nMUSE AGENT .MD & DRIVE AUDIT REPORT\n{'='*70}")
for acct, d in res.items():
if not d.get("connected"):
print(f"\n[AGENT {acct.upper()}] ✗ Connection failed: {d.get('error')}")
continue
score = d.get("drive_score", 0)
status_color = "HIGH DRIVE" if score >= 75 else ("PARTIAL" if score >= 50 else "LOW DRIVE / STALE")
print(f"\n[AGENT {acct.upper()}] DRIVE Score: {score}/100 ({status_color}) VM: {d.get('vm_id', 'unknown')}")
for fname, finfo in d.get("files", {}).items():
ex = "✓" if finfo.get("exists") else "✗"
sz = f"{finfo.get('size', 0):6} bytes"
mod = (finfo.get("modified") or "")[:19]
print(f" {ex} {fname:24} {sz} {mod}")
if d.get("issues"):
print(" Issues:")
for iss in d["issues"]:
print(f" • {iss}")
print(f"\n{'='*70}\n")
elif args.cmd == "list":
entries = list_files(args.account, args.path)
print(json.dumps(entries, indent=2))
elif args.cmd == "read":
r = read_md(args.account, args.filename)
print(r.get("text", ""))
elif args.cmd == "write":
content = args.content
if args.file:
content = Path(args.file).read_text(encoding="utf-8")
if content is None:
print("Error: provide --content or --file", file=sys.stderr)
sys.exit(2)
res = write_md(args.account, args.filename, content)
print(json.dumps(res, indent=2))
elif args.cmd == "diff":
res = diff_md(args.account, args.filename)
if res["identical"]:
print(f"{args.account}:{args.filename} matches local shared/operators/{args.filename} exactly.")
else:
print(res["diff"])
elif args.cmd == "amend":
content = args.content
if args.file:
content = Path(args.file).read_text(encoding="utf-8")
if content is None:
print("Error: provide --content or --file", file=sys.stderr)
sys.exit(2)
try:
res = amend_md(args.filename, content, author=args.author, reason=args.reason)
print(json.dumps(res, indent=2))
except Exception as e:
print(json.dumps({"ok": False, "error": str(e)}), indent=2)
sys.exit(1)
elif args.cmd == "append":
try:
res = append_md(args.filename, args.text, author=args.author, section=args.section)
print(json.dumps(res, indent=2))
except Exception as e:
print(json.dumps({"ok": False, "error": str(e)}), indent=2)
sys.exit(1)
elif args.cmd == "pull":
try:
res = pull_md(args.account, args.filename)
print(json.dumps(res, indent=2))
except Exception as e:
print(json.dumps({"ok": False, "error": str(e)}), indent=2)
sys.exit(1)
elif args.cmd == "inject-drive":
res = inject_drive(args.account, force=args.force)
print(json.dumps(res, indent=2))
elif args.cmd == "sync-all":
results = {}
for acct in VALID_ACCOUNTS:
try:
results[acct] = inject_drive(acct)
print(f"✓ Injected DRIVE into {acct}")
except Exception as e:
results[acct] = {"ok": False, "error": str(e)}
print(f"✗ Failed {acct}: {e}")
elif args.cmd == "ssh-info":
res = get_ssh_info(args.account)
print(json.dumps(res, indent=2))
if __name__ == "__main__":
main()
-1357
View File
File diff suppressed because it is too large Load Diff
-333
View File
@@ -1,333 +0,0 @@
#!/usr/bin/env python3
"""box-chat-cdp.py — READ-ONLY CDP scraper for box-chat.py.
Runs INSIDE the agent's netns (via netvm-exec.sh), where the agent's
headless Chromium CDP port is reachable on 127.0.0.1. Performs only
Runtime.evaluate reads plus benign navigation clicks (panel open, thread
switch, restore-to-main). Never sends messages, never touches the
composer, never clicks send.
Usage:
box-chat-cdp.py <agent> threads
box-chat-cdp.py <agent> messages <thread-id>
Prints exactly one JSON object to stdout. Exit 0 on success, 1 on
failure (stdout still carries {"ok": false, ...}).
"""
import importlib.util
import json
import sys
import time
import urllib.request
def load_accounts():
path = "/home/super/Projects/NetVM/bin/netvm-registry.py"
spec = importlib.util.spec_from_file_location("netvm_registry", path)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
accounts = {}
for node, rec in mod.load().items():
accounts[node] = (node, "http://127.0.0.1:%d/json/list" % rec["cdp_port"])
return accounts
def connect(agent):
accounts = load_accounts()
if agent not in accounts:
raise RuntimeError("unknown agent in registry: %s" % agent)
_node, cdp_url = accounts[agent]
with urllib.request.urlopen(cdp_url, timeout=10) as r:
targets = json.load(r)
pages = [t for t in targets if t.get("type") == "page"]
if not pages:
raise RuntimeError("no page target on CDP")
import websocket
ws = websocket.create_connection(
pages[0]["webSocketDebuggerUrl"], timeout=60)
return ws
def ev(ws, expr, await_p=False):
ws.send(json.dumps({
"id": 1, "method": "Runtime.evaluate",
"params": {"expression": expr, "returnByValue": True,
"awaitPromise": await_p},
}))
# Drain CDP events until we get our command response (id 1).
# The browser can emit events (Runtime.executionContextCreated, etc.)
# at any time; taking the first recv() blindly returns None on a
# busy page (observed as transient NO_SWITCHER / "unexpected
# messages payload" on pip/opm 2026-10-04).
resp = None
drained = 0
for _ in range(50):
raw = ws.recv()
resp = json.loads(raw)
if resp.get("id") == 1:
break
drained += 1
else:
raise RuntimeError("CDP: no response to Runtime.evaluate (drained %d)" % drained)
pass # drained count available in `drained` if needed
res = resp.get("result", {})
if res.get("subtype") == "error":
raise RuntimeError("JS error: %s" % str(res.get("description"))[:200])
return res.get("result", {}).get("value")
TITLE_OF = """const titleOf = r => {
const s = r.querySelector('span[title]');
return s ? s.getAttribute('title').trim()
: ((r.innerText||'').split('\\n')[0]||'').trim();
};"""
ENSURE = """(async () => {
const sleep = ms => new Promise(r => setTimeout(r, ms));
%s
if (window.location.pathname === '/thread/new') {
window.location.href = '/'; await sleep(4000);
}
const nav = document.querySelector('[data-testid="hatch-nav-chat"]');
if (!nav) return 'NO_CHAT_NAV';
if (nav.getAttribute('aria-current') !== 'page') { nav.click(); await sleep(3000); }
const panelOpen = () => !!document.querySelector('[data-testid="hatch-chat-compose"]');
if (!panelOpen()) {
// Retry: the switcher may not be rendered yet if the SPA is still
// settling (observed transient NO_SWITCHER on pip/opm 2026-10-04).
let sw = null;
for (let k = 0; k < 4 && !sw; k++) {
sw = document.querySelector('[data-testid="hatch-chat-switcher-trigger"]');
if (!sw) await sleep(2000);
}
if (!sw) return 'NO_SWITCHER';
sw.click(); await sleep(2500);
if (!panelOpen()) return 'PANEL_CLOSED';
}
return 'OK';
})()""" % TITLE_OF
def op_threads(ws):
st = ev(ws, ENSURE, await_p=True)
if st != "OK":
raise RuntimeError("could not reach chat panel: %s" % st)
js = """(async () => {
const sleep = ms => new Promise(r => setTimeout(r, ms));
%s
const snap = [...document.querySelectorAll('[data-testid="hatch-thread-row"]')]
.map(r => {
const spans = [...r.querySelectorAll('span')];
return {title: titleOf(r),
rel: spans.length ? (spans[spans.length-1].innerText||'').trim() : ''};
});
const ensurePanel = async () => {
if (document.querySelector('[data-testid="hatch-chat-compose"]')) return true;
for (let k = 0; k < 3; k++) {
const sw = document.querySelector('[data-testid="hatch-chat-switcher-trigger"]');
if (!sw) return false;
sw.click(); await sleep(3000);
if (document.querySelector('[data-testid="hatch-chat-compose"]')) return true;
}
return false;
};
const out = [];
for (const s of snap) {
const isMain = /^main chat$/i.test(s.title);
if (isMain) { out.push({id: 'main', kind: 'main', title: s.title, rel: s.rel}); continue; }
const panelOk = await ensurePanel();
if (!panelOk) { out.push({id: null, kind: 'sidechat', title: s.title, rel: s.rel, error: 'PANEL_CLOSED'}); continue; }
const row = [...document.querySelectorAll('[data-testid="hatch-thread-row"]')]
.find(r => titleOf(r) === s.title);
if (!row) { out.push({id: null, kind: 'sidechat', title: s.title, rel: s.rel, error: 'ROW_GONE'}); continue; }
row.click();
let id = null;
for (let t = 0; t < 20; t++) {
await sleep(500);
const m = window.location.href.match(/\\/thread\\/([0-9a-fA-F-]{36})/);
if (m) { id = m[1]; break; }
}
if (!id) {
const row2 = [...document.querySelectorAll('[data-testid="hatch-thread-row"]')]
.find(r => titleOf(r) === s.title);
if (row2) {
row2.click();
for (let t = 0; t < 12; t++) {
await sleep(500);
const m = window.location.href.match(/\\/thread\\/([0-9a-fA-F-]{36})/);
if (m) { id = m[1]; break; }
}
}
}
out.push({id, kind: 'sidechat', title: s.title, rel: s.rel});
}
return {threads: out};
})()""" % TITLE_OF
return ev(ws, js, await_p=True)
def op_messages(ws, thread_id):
st = ev(ws, ENSURE, await_p=True)
if st != "OK":
raise RuntimeError("could not reach chat panel: %s" % st)
js = """(async () => {
const sleep = ms => new Promise(r => setTimeout(r, ms));
%s
const THREAD = %s;
const ensurePanel = async () => {
if (document.querySelector('[data-testid="hatch-chat-compose"]')) return true;
for (let k = 0; k < 3; k++) {
const sw = document.querySelector('[data-testid="hatch-chat-switcher-trigger"]');
if (!sw) return false;
sw.click(); await sleep(3000);
if (document.querySelector('[data-testid="hatch-chat-compose"]')) return true;
}
return false;
};
const uuidOf = () => {
const m = window.location.href.match(/\\/thread\\/([0-9a-fA-F-]{36})/);
return m ? m[1].toLowerCase() : null;
};
let landed = false;
if (THREAD === 'main') {
await ensurePanel();
const row = [...document.querySelectorAll('[data-testid="hatch-thread-row"]')]
.find(r => /^main chat$/i.test(titleOf(r)));
if (row) { row.click(); await sleep(2500); }
landed = /muse\\.ai\\/?$/.test(window.location.href) && !uuidOf();
} else {
// already there?
if (uuidOf() === THREAD.toLowerCase()) landed = true;
// click rows until the URL carries our thread uuid (SPA navigation,
// keeps the JS context alive unlike location.href assignment)
for (let i = 0; i < 12 && !landed; i++) {
await ensurePanel();
const rows = [...document.querySelectorAll('[data-testid="hatch-thread-row"]')];
if (!rows.length) break;
const row = rows[i %% rows.length];
row.click();
for (let t = 0; t < 14; t++) {
await sleep(500);
if (uuidOf() === THREAD.toLowerCase()) { landed = true; break; }
if (uuidOf()) break; // navigated somewhere else; try next row
}
}
if (!landed) {
// fallback: SPA history navigation (no full page load)
window.history.pushState({}, '', '/thread/' + THREAD);
window.dispatchEvent(new PopStateEvent('popstate'));
await sleep(5000);
landed = uuidOf() === THREAD.toLowerCase();
}
}
if (!landed) return {error: 'THREAD_NOT_FOUND'};
await sleep(2500);
const sc = document.getElementById('hatch-chat-scroll');
if (sc) {
for (let i = 0; i < 3; i++) { sc.scrollTop = 0; await sleep(1500); }
sc.scrollTop = sc.scrollHeight; await sleep(800);
}
const els = [...document.querySelectorAll('[data-message-id]')];
const messages = els.map(m => {
const id = m.getAttribute('data-message-id');
const ps = [...m.querySelectorAll('p')].map(p => (p.innerText||'').trim()).filter(Boolean);
let text = ps.join('\\n');
if (!text) text = (m.innerText||'').replace(/^(Assistant message:|User message:)\\s*/, '').trim();
const t = m.querySelector('time');
return {id,
author: id.indexOf('assistant-msg') === 0 ? 'assistant' : 'user',
text,
ts: t ? (t.getAttribute('datetime') || t.innerText || null) : null};
});
return {url: window.location.href, messages};
})()""" % (TITLE_OF, json.dumps(thread_id))
data = ev(ws, js, await_p=True)
if not isinstance(data, dict) or "messages" not in data:
if isinstance(data, dict) and data.get("error") == "THREAD_NOT_FOUND":
raise RuntimeError("THREAD_NOT_FOUND: no such thread for this agent")
raise RuntimeError("unexpected messages payload")
return data
def restore_main(ws):
try:
ev(ws, """(() => {
%s
const row = [...document.querySelectorAll('[data-testid="hatch-thread-row"]')]
.find(r => /^main chat$/i.test(titleOf(r)));
if (row) row.click();
return 'ok';
})()""" % TITLE_OF)
except Exception:
pass
def main(argv):
if len(argv) < 3:
print(json.dumps({"ok": False, "code": "BAD_ARGS",
"error": "usage: box-chat-cdp.py <agent> threads|messages [thread-id]"}))
return 1
agent, op = argv[1], argv[2]
ws = None
try:
# CDP evaluate can resolve null when the page is mid-navigation
# (agent actively using browser). Reconnect fresh on each retry
# so we attach to the current page, not a stale JS context.
# Observed 2026-10-04: opm's browser navigates during reads.
data, last_err = None, None
for attempt in range(3):
try:
if ws is not None:
try:
ws.close()
except Exception:
pass
ws = connect(agent)
if op == "threads":
data = op_threads(ws)
ok = isinstance(data, dict) and isinstance(
data.get("threads"), list)
elif op == "messages":
if len(argv) < 4:
raise RuntimeError("messages requires thread-id")
data = op_messages(ws, argv[3])
ok = isinstance(data, dict) and isinstance(
data.get("messages"), list)
else:
raise RuntimeError("unknown op: %s" % op)
if ok:
break
last_err = "empty CDP result"
data = None
except RuntimeError as e:
last_err = str(e)
data = None
time.sleep(3)
if data is None:
raise RuntimeError(last_err or "CDP returned no usable data")
if op == "threads":
print(json.dumps({"ok": True, "agent": agent,
"threads": data["threads"]}))
else:
print(json.dumps({"ok": True, "agent": agent, "url": data["url"],
"messages": data["messages"]}))
return 0
except Exception as e:
print(json.dumps({"ok": False, "code": "CDP_ERROR",
"error": str(e)[:300]}))
return 1
finally:
if ws is not None:
try:
restore_main(ws)
except Exception:
pass
try:
ws.close()
except Exception:
pass
if __name__ == "__main__":
sys.exit(main(sys.argv))
-380
View File
@@ -1,380 +0,0 @@
#!/usr/bin/env python3
"""box-chat.py — allowlisted bl helper for Box thread oversight (READ-ONLY).
The board server (VM) never scrapes browsers directly. All thread reads go
through this helper, invoked as:
/home/super/Projects/NetVM/bin/box-chat.py thread-list <agent> [--kind K] [--limit N]
/home/super/Projects/NetVM/bin/box-chat.py thread-messages <agent> <thread-id> [--limit N] [--before MSGID]
Security properties (mirrors box-ctl.py):
- Fixed verb set; every argument validated before acting.
- <agent> must be a known node (muse, pip, 646, opm); anything else exits
before any netns/SSH/CDP work.
- <thread-id> must match ^[a-zA-Z0-9-]{1,64}$ or be the literal "main".
- READ-ONLY by construction: the CDP driver (box-chat-cdp.py) only runs
Runtime.evaluate reads plus benign navigation clicks. No sends, no
composer interaction, no shell=True anywhere. All subprocess calls use
argv lists.
- Every invocation audit-logged to box-chat.jsonl with caller identity.
Output: JSON to stdout ({"ok": true, ...} or {"ok": false, ...}),
exit 0 on success, nonzero on failure.
"""
import json
import os
import re
import subprocess
import sys
from datetime import datetime, timedelta, timezone
from pathlib import Path
NETVM_ROOT = Path("/home/super/Projects/NetVM")
BIN = NETVM_ROOT / "bin"
CDP_DRIVER = BIN / "box-chat-cdp.py"
NETVM_EXEC = BIN / "netvm-exec.sh"
CHAT_LOG = NETVM_ROOT / "box-chat.jsonl"
DM_LOG = NETVM_ROOT / "dm-log.jsonl"
VALID_AGENTS = {"muse", "pip", "646", "opm"}
VALID_KINDS = {"main", "sidechat", "dm", "all"}
AGENT_RE = re.compile(r"^[a-z0-9-]{1,64}$")
THREAD_RE = re.compile(r"^[a-zA-Z0-9-]{1,64}$")
MSGID_RE = re.compile(r"^[a-zA-Z0-9-]{1,128}$")
REL_MONTHS = {m: i + 1 for i, m in enumerate(
["jan", "feb", "mar", "apr", "may", "jun",
"jul", "aug", "sep", "oct", "nov", "dec"])}
def utcnow():
return datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
def out(ok, **kw):
payload = {"ok": ok}
payload.update(kw)
print(json.dumps(payload))
def fail(code, error, detail=None, exit_code=1):
payload = {"ok": False, "code": code, "error": error}
if detail is not None:
payload["detail"] = detail
print(json.dumps(payload))
sys.exit(exit_code)
def audit(action, agent=None, name=None, kind=None):
"""Append invocation record to the bl-side log."""
try:
entry = {
"ts": utcnow(),
"action": action,
"agent": agent,
"name": name,
"kind": kind,
"caller": os.environ.get("BOX_CALLER", "unknown"),
}
with open(CHAT_LOG, "a") as f:
f.write(json.dumps(entry) + "\n")
except Exception:
pass # audit failure must not break the action
def check_agent(agent):
if not agent or agent not in VALID_AGENTS:
fail("BAD_AGENT", "agent must be one of %s" % sorted(VALID_AGENTS),
{"field": "agent", "value": agent})
return agent
def check_thread_id(tid):
if tid != "main" and not (tid and THREAD_RE.match(tid)):
fail("BAD_THREAD", "thread-id must be 'main' or match ^[a-zA-Z0-9-]{1,64}$",
{"field": "thread-id", "value": tid})
return tid
def check_limit(val, default):
if val is None:
return default
try:
n = int(val)
except (TypeError, ValueError):
fail("BAD_LIMIT", "limit must be an integer 1-200", {"value": val})
if not 1 <= n <= 200:
fail("BAD_LIMIT", "limit must be an integer 1-200", {"value": val})
return n
def run_cdp(agent, op, args, timeout):
"""Run the CDP driver inside the agent's netns. argv only, no shell."""
cmd = [str(NETVM_EXEC), agent, "--", sys.executable,
str(CDP_DRIVER), agent, op] + args
try:
r = subprocess.run(cmd, capture_output=True, text=True,
timeout=timeout)
except subprocess.TimeoutExpired:
fail("CDP_TIMEOUT", "CDP driver timed out in netns",
{"agent": agent, "op": op})
if r.returncode != 0 and not r.stdout.strip():
fail("CDP_ERROR", "CDP driver failed",
{"stderr": (r.stderr or "")[-300:]})
try:
data = json.loads(r.stdout.strip())
except Exception:
fail("CDP_ERROR", "CDP driver returned non-JSON",
{"stdout": r.stdout[-300:], "stderr": (r.stderr or "")[-300:]})
if not data.get("ok"):
code = data.get("code", "CDP_ERROR")
if "THREAD_NOT_FOUND" in str(data.get("error", "")):
code = "THREAD_NOT_FOUND"
fail(code, data.get("error", "cdp failed"))
return data
def parse_rel_time(rel):
"""Panel relative times ('just now', '4m', '2h', '3d', 'Oct 3') ->
(approx ISO8601, True). Returns (None, False) when unparseable."""
now = datetime.now(timezone.utc)
rel = (rel or "").strip().lower()
if rel in ("just now", "now"):
return now.strftime("%Y-%m-%dT%H:%M:%SZ"), True
m = re.match(r"^(\d+)\s*m(in(ute)?s?)?$", rel)
if m:
return (now - timedelta(minutes=int(m.group(1)))).strftime("%Y-%m-%dT%H:%M:%SZ"), True
m = re.match(r"^(\d+)\s*h((ou)?rs?)?$", rel)
if m:
return (now - timedelta(hours=int(m.group(1)))).strftime("%Y-%m-%dT%H:%M:%SZ"), True
m = re.match(r"^(\d+)\s*d(ays?)?$", rel)
if m:
return (now - timedelta(days=int(m.group(1)))).strftime("%Y-%m-%dT%H:%M:%SZ"), True
m = re.match(r"^([a-z]{3})\s+(\d{1,2})$", rel)
if m and m.group(1) in REL_MONTHS:
dt = now.replace(month=REL_MONTHS[m.group(1)], day=int(m.group(2)),
hour=12, minute=0, second=0, microsecond=0)
if dt > now:
dt = dt.replace(year=dt.year - 1)
return dt.strftime("%Y-%m-%dT%H:%M:%SZ"), True
return None, False
def dm_conversations(agent):
"""DM conversations involving <agent>, synthesized from dm-log.jsonl.
Real timestamps and counts; bodies are not stored in the log (by design)
— message text for these threads is the wire tag, and full bodies live
in the target chat's messages.
"""
convos = {}
try:
with open(DM_LOG) as f:
for line in f:
line = line.strip()
if not line:
continue
try:
e = json.loads(line)
except Exception:
continue
if e.get("type") not in ("sent", "send_done"):
continue
frm, to = e.get("agent"), e.get("to")
if not frm or not to:
continue
if agent not in (frm, to):
continue
key = tuple(sorted([frm, to]))
c = convos.setdefault(key, {"count": 0, "last_ts": "",
"entries": []})
c["count"] += 1
if e.get("ts", "") > c["last_ts"]:
c["last_ts"] = e["ts"]
c["entries"].append(e)
except FileNotFoundError:
return []
out = []
for (a, b), c in sorted(convos.items()):
out.append({
"id": "dm-%s-%s" % (a, b),
"kind": "dm",
"title": "dm:%s:%s" % (a, b),
"participants": [a, b],
"last_message_at": c["last_ts"] or None,
"last_message_approx": False,
"message_count": c["count"],
})
return out
def dm_thread_messages(agent, thread_id):
"""Messages for a dm-<a>-<b> thread, from dm-log.jsonl (complete log)."""
parts = thread_id[3:].split("-")
if len(parts) != 2:
fail("THREAD_NOT_FOUND", "no such DM thread: %s" % thread_id)
a, b = parts
if agent not in (a, b):
fail("THREAD_NOT_FOUND", "no such DM thread: %s" % thread_id)
msgs = []
try:
with open(DM_LOG) as f:
for line in f:
line = line.strip()
if not line:
continue
try:
e = json.loads(line)
except Exception:
continue
if e.get("type") not in ("sent", "send_done"):
continue
frm, to = e.get("agent"), e.get("to")
if not frm or not to:
continue
if tuple(sorted([frm, to])) != (a, b):
continue
sender = e.get("agent")
msgs.append({
"id": str(e.get("id", "")),
"from": {"role": "agent", "name": sender},
"text": "[from:%s] [id:%s] -> %s" % (
sender, e.get("id"), e.get("target", "?")),
"ts": e.get("ts"),
"target": e.get("target"),
"verified": e.get("verified"),
})
except FileNotFoundError:
pass
msgs.sort(key=lambda m: m.get("ts") or "")
# de-dupe send_done/sent pairs on DM id, keep the richer record
seen = {}
for m in msgs:
prev = seen.get(m["id"])
if prev is None or (m.get("verified") and not prev.get("verified")):
seen[m["id"]] = m
msgs = sorted(seen.values(), key=lambda m: m.get("ts") or "")
return msgs
def act_thread_list(agent, kind, limit):
threads = []
if kind in ("main", "sidechat", "all"):
data = run_cdp(agent, "threads", [], timeout=240)
for t in data.get("threads", [])[:limit]:
if kind != "all" and t.get("kind") != kind:
continue
ts, approx = parse_rel_time(t.get("rel", ""))
threads.append({
"id": t.get("id"),
"kind": t.get("kind"),
"title": t.get("title"),
"participants": [agent, "human"],
"last_message_at": ts,
"last_message_approx": approx,
"message_count": None, # list is a panel scrape; counts need a thread open
})
if kind in ("dm", "all"):
threads.extend(dm_conversations(agent))
audit("thread-list", agent=agent, kind=kind)
out(True, agent=agent, kind=kind, threads=threads,
fetched_at=utcnow())
def act_thread_messages(agent, thread_id, limit, before):
if thread_id.startswith("dm-"):
msgs = dm_thread_messages(agent, thread_id)
thread = {"id": thread_id, "kind": "dm",
"title": "dm:%s" % thread_id[3:].replace("-", ":"),
"participants": thread_id[3:].split("-")}
else:
data = run_cdp(agent, "messages", [thread_id], timeout=150)
raw = data.get("messages", [])
thread = {"id": thread_id,
"kind": "main" if thread_id == "main" else "sidechat",
"participants": [agent, "human"]}
msgs = [{
"id": m.get("id"),
"from": ({"role": "agent", "name": agent}
if m.get("author") == "assistant"
else {"role": "human", "name": "human"}),
"text": m.get("text", ""),
"ts": m.get("ts"),
} for m in raw]
if before:
if not MSGID_RE.match(before):
fail("BAD_CURSOR", "before must match ^[a-zA-Z0-9-]{1,128}$",
{"value": before})
idx = next((i for i, m in enumerate(msgs) if m["id"] == before), None)
if idx is None:
fail("BAD_CURSOR", "no message with that id in loaded window",
{"value": before})
msgs = msgs[:idx]
older = len(msgs)
if len(msgs) > limit:
msgs = msgs[-limit:]
next_before = msgs[0]["id"] if older > len(msgs) and msgs else None
audit("thread-messages", agent=agent, name=thread_id)
out(True, agent=agent, thread=thread, messages=msgs,
next_before=next_before, loaded_count=older, fetched_at=utcnow())
USAGE = """usage: box-chat.py <action> [args]
thread-list <agent> [--kind main|sidechat|dm|all] [--limit N]
thread-messages <agent> <thread-id> [--limit N] [--before MSGID]
read-only. <agent> is one of: muse, pip, 646, opm."""
def parse_flags(rest, names):
"""Parse [--flag value] pairs; returns (positionals, {flag: value})."""
pos, flags = [], {}
i = 0
while i < len(rest):
tok = rest[i]
if tok.startswith("--") and tok[2:] in names:
if i + 1 >= len(rest):
fail("BAD_ARGS", "flag %s needs a value" % tok)
flags[tok[2:]] = rest[i + 1]
i += 2
elif tok.startswith("--"):
fail("BAD_ARGS", "unknown flag: %s" % tok)
else:
pos.append(tok)
i += 1
return pos, flags
def main(argv):
if len(argv) < 2:
print(USAGE, file=sys.stderr)
sys.exit(2)
action = argv[1]
if action == "thread-list":
pos, flags = parse_flags(argv[2:], {"kind", "limit"})
if len(pos) != 1:
fail("BAD_ARGS", "usage: thread-list <agent> [--kind K] [--limit N]")
agent = check_agent(pos[0])
kind = flags.get("kind", "all")
if kind not in VALID_KINDS:
fail("BAD_ARGS", "kind must be one of %s" % sorted(VALID_KINDS))
act_thread_list(agent, kind, check_limit(flags.get("limit"), 50))
elif action == "thread-messages":
pos, flags = parse_flags(argv[2:], {"limit", "before"})
if len(pos) != 2:
fail("BAD_ARGS",
"usage: thread-messages <agent> <thread-id> [--limit N] [--before MSGID]")
agent = check_agent(pos[0])
tid = check_thread_id(pos[1])
act_thread_messages(agent, tid, check_limit(flags.get("limit"), 50),
flags.get("before"))
else:
print(USAGE, file=sys.stderr)
fail("BAD_ARGS", "unknown action: %s" % action)
if __name__ == "__main__":
main(sys.argv)
-4763
View File
File diff suppressed because it is too large Load Diff
+1611
View File
File diff suppressed because it is too large Load Diff
-720
View File
@@ -1,720 +0,0 @@
#!/usr/bin/env python3
"""box-onboard-tui.py — Dedicated interactive TUI for NetVM Onboard Connects, Tmux Workers, and Auto-Approvals.
Features 4 bridged surfaces:
[1 / F1] Onboard Connects:
Live inventory of fleet nodes & client onboarding pipelines,
invite codes, token feeding urgency, OTP verification, and salvage dispatch.
[2 / F2] Tmux Workers & Tally:
Multi-socket worker inventory (/tmp/tmux-1000/default, lte, muse.sock, netns socks),
active panes, live pane scrollback preview, worker spawning, and session killing.
[3 / F3] Auto-Approvals & Regex Matcher:
Master auto-approval toggle, per-agent policies, terminal regex rule engine
(Muse Code runs, A/B/C choices, 1/2 menus, y/n confirmations), and interactive regex tester.
[4 / F4] Box Surface & Logs:
Surface link with https://box.muse-dev.online/, real-time audit log stream,
and search/filter capabilities.
"""
from __future__ import annotations
import curses
import json
import os
import re
import subprocess
import sys
import time
from dataclasses import asdict
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple
REPO_ROOT = Path(__file__).resolve().parent.parent
BIN_DIR = REPO_ROOT / "bin"
sys.path.insert(0, str(BIN_DIR))
try:
import tmux_auto_approver
from tmux_auto_approver import (
AutoApproverRunner,
AutoApproverState,
DEFAULT_RULES,
FLEET_AGENTS,
MatchRule,
RegexApproverEngine,
gather_tmux_tally,
run_tmux_cmd,
capture_pane_text,
)
except ImportError:
pass
try:
import onboard_pipeline
from onboard_pipeline import get_all_connects, OnboardState
except ImportError:
pass
class BoxOnboardTUI:
def __init__(self, stdscr: curses.window):
self.stdscr = stdscr
self.current_tab = 0 # 0: Onboard, 1: Tmux, 2: Auto-Approvals, 3: Surface & Logs
self.tabs = [
"1: ONBOARD CONNECTS",
"2: TMUX WORKERS & TALLY",
"3: AUTO-APPROVALS & REGEX",
"4: SURFACE & LOGS",
]
# Curses initialization
try:
curses.curs_set(0)
except Exception:
pass
self.stdscr.nodelay(True)
self.stdscr.keypad(True)
if hasattr(curses, "set_escdelay"):
try:
curses.set_escdelay(25)
except Exception:
pass
try:
curses.mousemask(curses.ALL_MOUSE_EVENTS | curses.REPORT_MOUSE_POSITION)
if hasattr(curses, "mouseinterval"):
curses.mouseinterval(0)
sys.stdout.write("\033[?1000h\033[?1002h\033[?1006h\033[?2004h")
sys.stdout.flush()
except Exception:
pass
self._init_colors()
# Shared State
self.approver_state = AutoApproverState.load()
self.tally = gather_tmux_tally(self.approver_state)
self.connects: List[Dict[str, Any]] = []
self._refresh_connects()
# Selection indices
self.sel_connect_idx = 0
self.sel_pane_idx = 0
self.sel_rule_idx = 0
self.sel_log_scroll = 0
# UI State & Modals
self.modal: Optional[str] = None # "spawn_worker", "submit_otp", "test_regex", "help"
self.modal_input_buf = ""
self.modal_input_cursor = 0
self.toast_msg = "Welcome to Box Onboard & Tmux Console. Press ? for help."
self.toast_level = "info"
self.toast_time = time.time() + 4.0
# Interactive Regex Matcher state (Tab 3)
self.test_text_buf = "Would you like to run the following\n\n $ box fleet status\n\n› 1. Yes, proceed (y)\n 2. No, and tell Muse Code what to do"
self.test_match_verdict: Optional[Dict[str, Any]] = None
self._eval_test_match()
# Runner instance for one-shot runs
self.runner = AutoApproverRunner()
self.last_auto_poll = 0.0
def _init_colors(self) -> None:
try:
curses.start_color()
curses.use_default_colors()
curses.init_pair(1, curses.COLOR_CYAN, -1) # Accent / Info
curses.init_pair(2, curses.COLOR_YELLOW, -1) # Warning / Agent
curses.init_pair(3, curses.COLOR_GREEN, -1) # Success / Active
curses.init_pair(4, curses.COLOR_RED, -1) # Error / Blocked
curses.init_pair(5, curses.COLOR_MAGENTA, -1) # Special / Category
curses.init_pair(6, curses.COLOR_BLACK, curses.COLOR_CYAN) # Header Selected
curses.init_pair(7, curses.COLOR_BLACK, curses.COLOR_WHITE) # Selected Row
curses.init_pair(8, curses.COLOR_BLACK, curses.COLOR_YELLOW) # Warning Banner
curses.init_pair(9, curses.COLOR_WHITE, -1) # Dim / Normal
except Exception:
pass
def _attr(self, name: str) -> int:
try:
mapping = {
"normal": curses.color_pair(0),
"cyan": curses.color_pair(1) | curses.A_BOLD,
"yellow": curses.color_pair(2) | curses.A_BOLD,
"green": curses.color_pair(3) | curses.A_BOLD,
"red": curses.color_pair(4) | curses.A_BOLD,
"magenta": curses.color_pair(5) | curses.A_BOLD,
"head_sel": curses.color_pair(6) | curses.A_BOLD,
"row_sel": curses.color_pair(7) | curses.A_BOLD,
"warn_banner": curses.color_pair(8) | curses.A_BOLD,
"dim": curses.color_pair(9) | curses.A_DIM,
}
return mapping.get(name, 0)
except Exception:
return 0
def set_toast(self, msg: str, level: str = "info") -> None:
self.toast_msg = msg
self.toast_level = level
self.toast_time = time.time() + 4.0
def _refresh_connects(self) -> None:
try:
self.connects = get_all_connects(fast=True)
except Exception:
self.connects = []
def _refresh_tally(self) -> None:
self.approver_state = AutoApproverState.load()
self.tally = gather_tmux_tally(self.approver_state)
def _eval_test_match(self) -> None:
engine = RegexApproverEngine([MatchRule(**r) for r in self.approver_state.rules])
v = engine.evaluate(self.test_text_buf)
self.test_match_verdict = {
"matched": v.matched,
"rule_name": v.rule_name,
"key": v.key,
"category": v.category,
"press_enter": v.press_enter,
"reason": v.reason,
"is_blocked": v.is_blocked,
"blocked_reason": v.blocked_reason,
"excerpt": v.excerpt,
}
# =================================================================
# Render Helpers
# =================================================================
def safe_addstr(self, y: int, x: int, text: str, attr: int = 0) -> None:
h, w = self.stdscr.getmaxyx()
if 0 <= y < h and 0 <= x < w:
avail = max(0, w - x - 1)
try:
self.stdscr.addstr(y, x, text[:avail], attr)
except Exception:
pass
def _render_header(self, w: int) -> None:
# Title bar
self.safe_addstr(0, 0, " " * w, self._attr("header_sel"))
title = " 󰢹 BOX ONBOARD & TMUX CONSOLE [https://box.muse-dev.online/] "
self.safe_addstr(0, 1, title, self._attr("header_sel"))
# Master auto-approve badge in header
master_tag = " [● AUTO-APPROVE: ON] " if self.approver_state.global_enabled else " [○ AUTO-APPROVE: OFF] "
m_attr = self._attr("green") if self.approver_state.global_enabled else self._attr("warn_banner")
self.safe_addstr(0, max(len(title) + 2, w - len(master_tag) - 2), master_tag, m_attr)
# Tab navigation bar
self.safe_addstr(1, 0, " " * w, self._attr("dim"))
col = 1
for idx, tab_name in enumerate(self.tabs):
is_cur = (idx == self.current_tab)
pill = f" [{tab_name}] "
attr = self._attr("header_sel") if is_cur else self._attr("dim")
self.safe_addstr(1, col, pill, attr)
col += len(pill) + 2
self.safe_addstr(2, 0, "─" * w, self._attr("dim"))
def _render_footer(self, h: int, w: int) -> None:
y = h - 2
self.safe_addstr(y, 0, "─" * w, self._attr("dim"))
# Toast or Hints
if time.time() < self.toast_time:
attr = self._attr("green") if self.toast_level == "success" else (self._attr("warn_banner") if self.toast_level == "warn" else self._attr("cyan"))
self.safe_addstr(y + 1, 1, f" 󰋼 {self.toast_msg} ", attr)
else:
if self.current_tab == 0:
hints = " [ONBOARD] 1-4: Tabs j/k: Select n: New Onboard o: Submit OTP s: Salvage WO r: Refresh ?: Help q: Quit"
elif self.current_tab == 1:
hints = " [TMUX] 1-4: Tabs j/k: Select t: Toggle Auto-Approve n: Spawn k: Kill p: Prune Enter: Full Tail q: Quit"
elif self.current_tab == 2:
hints = " [RULES] 1-4: Tabs Space/a: Toggle Master m: Test Matcher j/k: Rules o: Trigger Once q: Quit"
else:
hints = " [LOGS] 1-4: Tabs j/k: Scroll e: Sync Box API r: Refresh q: Quit"
self.safe_addstr(y + 1, 1, hints[:w - 2], self._attr("dim"))
# =================================================================
# Tab 1: Onboard Connects
# =================================================================
def _render_tab_onboard(self, h: int, w: int) -> None:
start_y = 3
max_rows = h - 7
self.safe_addstr(start_y, 2, "ACTIVE FLEET AGENTS & CLIENT ONBOARDING CONNECTS", self._attr("cyan"))
self.safe_addstr(start_y + 1, 2, "─" * (w - 4), self._attr("dim"))
hdr = f" {'NODE':<8} {'TYPE':<16} {'STAGE / STATUS':<18} {'CDP':<8} {'INVITE':<10} {'ROLE / DETAIL'}"
self.safe_addstr(start_y + 2, 2, hdr, self._attr("bold"))
self.safe_addstr(start_y + 3, 2, "─" * (w - 4), self._attr("dim"))
if not self.connects:
self.safe_addstr(start_y + 4, 4, "(No onboard records found. Press 'n' to initiate client onboarding)", self._attr("dim"))
return
for idx, c in enumerate(self.connects[:max_rows]):
row_y = start_y + 4 + idx
is_sel = (idx == self.sel_connect_idx)
node = c.get("node", "")
t_str = c.get("type", "")
st_str = c.get("stage", c.get("status", ""))
cdp = str(c.get("cdp_port") or "-")
code = c.get("invite_code") or "-"
role = c.get("role") or c.get("detail") or c.get("email") or ""
line = f" {node:<8} {t_str:<16} {st_str:<18} {cdp:<8} {code:<10} {role}"
attr = self._attr("selected") if is_sel else (self._attr("green") if "active" in st_str or "completed" in st_str else self._attr("normal"))
self.safe_addstr(row_y, 2, " " * (w - 4), attr if is_sel else 0)
self.safe_addstr(row_y, 2, line, attr)
# =================================================================
# Tab 2: Tmux Workers & Tally
# =================================================================
def _render_tab_tmux(self, h: int, w: int) -> None:
start_y = 3
split_h = max(6, (h - 6) // 2)
# Header summary
summary = f"TMUX WORKER TALLY — {self.tally.total_sessions} Sessions · {self.tally.total_panes} Panes · {self.tally.active_workers} Active across {self.tally.total_sockets} Sockets"
self.safe_addstr(start_y, 2, summary, self._attr("cyan"))
hdr = f" {'Socket':<18} {'Session':<14} {'Pane':<6} {'PID':<8} {'Agent':<6} {'Cmd':<16} {'Auto-Approve'}"
self.safe_addstr(start_y + 1, 2, hdr, self._attr("bold"))
self.safe_addstr(start_y + 2, 2, "─" * (w - 4), self._attr("dim"))
panes = self.tally.panes
table_rows = split_h - 3
if not panes:
self.safe_addstr(start_y + 3, 4, "(No active tmux sessions found)", self._attr("dim"))
else:
for idx, p in enumerate(panes[:table_rows]):
row_y = start_y + 3 + idx
is_sel = (idx == self.sel_pane_idx)
sock_short = os.path.basename(p.socket)
auto_str = "[AUTO: ON]" if p.auto_approve else "[AUTO: OFF]"
auto_attr = self._attr("green") if p.auto_approve else self._attr("dim")
line = f" {sock_short:<18} {p.session[:13]:<14} {p.pane_id:<6} {p.pane_pid:<8} {p.agent_node:<6} {p.current_command[:15]:<16} {auto_str}"
attr = self._attr("selected") if is_sel else self._attr("normal")
self.safe_addstr(row_y, 2, " " * (w - 4), attr if is_sel else 0)
self.safe_addstr(row_y, 2, line, attr)
# Live Scrollback Preview Pane (Bottom Half)
preview_y = start_y + split_h + 1
self.safe_addstr(preview_y - 1, 2, "─" * (w - 4), self._attr("dim"))
sel_pane = panes[self.sel_pane_idx] if (panes and 0 <= self.sel_pane_idx < len(panes)) else None
if sel_pane:
prev_hdr = f"LIVE SCROLLBACK PREVIEW — {sel_pane.session}:{sel_pane.pane_id} ({sel_pane.current_command}) on {sel_pane.socket}"
self.safe_addstr(preview_y, 2, prev_hdr, self._attr("yellow"))
txt = capture_pane_text(sel_pane.socket, sel_pane.pane_id, lines=h - preview_y - 4)
p_lines = txt.strip().splitlines()
for r_i, l_str in enumerate(p_lines[:h - preview_y - 4]):
self.safe_addstr(preview_y + 1 + r_i, 3, l_str, self._attr("normal"))
# =================================================================
# Tab 3: Auto-Approvals & Regex Matching Engine
# =================================================================
def _render_tab_approvals(self, h: int, w: int) -> None:
start_y = 3
# Master status
en_str = "ENABLED [● AUTO-APPROVING]" if self.approver_state.global_enabled else "DISABLED [○ MANUAL APPROVALS ONLY]"
en_attr = self._attr("green") if self.approver_state.global_enabled else self._attr("warn_banner")
self.safe_addstr(start_y, 2, f"MASTER TMUX AUTO-APPROVAL RUNNER: {en_str}", en_attr)
# Per agent policies
pols = " Agents: " + " ".join([f"{a}: {'ON [✔]' if self.approver_state.agents_enabled.get(a, True) else 'OFF [✖]'}" for a in FLEET_AGENTS])
self.safe_addstr(start_y + 1, 2, pols, self._attr("dim"))
self.safe_addstr(start_y + 2, 2, "─" * (w - 4), self._attr("dim"))
# Rules Table
self.safe_addstr(start_y + 3, 2, "ACTIVE TERMINAL REGEX APPROVAL RULES:", self._attr("cyan"))
hdr = f" {'STATUS':<8} {'RULE ID':<26} {'CATEGORY':<14} {'KEY':<6} {'DESCRIPTION'}"
self.safe_addstr(start_y + 4, 2, hdr, self._attr("bold"))
self.safe_addstr(start_y + 5, 2, "─" * (w - 4), self._attr("dim"))
rules = self.approver_state.rules
for idx, r in enumerate(rules[:6]):
row_y = start_y + 6 + idx
is_sel = (idx == self.sel_rule_idx)
st_tag = "ACTIVE" if r.get("enabled") else "OFF"
line = f" {st_tag:<8} {r.get('id'):<26} {r.get('category'):<14} {r.get('response_key'):<6} {r.get('description', '')[:35]}"
attr = self._attr("selected") if is_sel else self._attr("normal")
self.safe_addstr(row_y, 2, " " * (w - 4), attr if is_sel else 0)
self.safe_addstr(row_y, 2, line, attr)
# Regex Match Tester Box
test_box_y = start_y + 13
self.safe_addstr(test_box_y, 2, "─" * (w - 4), self._attr("dim"))
self.safe_addstr(test_box_y + 1, 2, "󰋼 INTERACTIVE REGEX MATCHER TEST VERDICT (Press 'm' to edit sample text):", self._attr("yellow"))
if self.test_match_verdict:
matched = self.test_match_verdict.get("matched")
if matched:
r_name = self.test_match_verdict.get("rule_name")
k = self.test_match_verdict.get("key")
res_str = f"✔ MATCHED: Rule '{r_name}' -> Auto-Replies: '{k}'"
self.safe_addstr(test_box_y + 2, 4, res_str, self._attr("green"))
elif self.test_match_verdict.get("is_blocked"):
b_reason = self.test_match_verdict.get("blocked_reason")
self.safe_addstr(test_box_y + 2, 4, f"✖ BLOCKED: {b_reason}", self._attr("red"))
else:
self.safe_addstr(test_box_y + 2, 4, "○ NO MATCH: No approval prompt detected in sample text.", self._attr("dim"))
# Excerpt
exc = self.test_match_verdict.get("excerpt") or ""
if exc:
self.safe_addstr(test_box_y + 3, 4, f"Matched Excerpt: '{exc.strip().replace(chr(10), ' ')[:70]}'", self._attr("cyan"))
# =================================================================
# Tab 4: Surface & Logs (https://box.muse-dev.online/)
# =================================================================
def _render_tab_logs(self, h: int, w: int) -> None:
start_y = 3
self.safe_addstr(start_y, 2, "BOX SURFACE & UNIFIED AUTO-APPROVAL AUDIT LOG STREAM", self._attr("cyan"))
self.safe_addstr(start_y + 1, 2, "Surface link: https://box.muse-dev.online/ · HTTPS Exec: https://exec.muse-dev.online/exec", self._attr("dim"))
self.safe_addstr(start_y + 2, 2, "─" * (w - 4), self._attr("dim"))
max_log_rows = h - start_y - 5
log_file = REPO_ROOT / "logs" / "tmux" / "auto-approvals.jsonl"
lines = []
if log_file.exists():
try:
with open(log_file) as f:
lines = f.readlines()
except Exception:
pass
if not lines:
self.safe_addstr(start_y + 4, 4, "(No auto-approval log events recorded yet. Press 'o' on Rules tab to run once)", self._attr("dim"))
return
tail = lines[-(max_log_rows + self.sel_log_scroll):]
if self.sel_log_scroll > 0:
tail = tail[:-self.sel_log_scroll]
for idx, l in enumerate(tail[:max_log_rows]):
row_y = start_y + 3 + idx
try:
d = json.loads(l)
ts = d.get("timestamp", "")[:19].replace("T", " ")
act = d.get("action", "")
ag = d.get("agent", "")
sess = d.get("session", "")
pane = d.get("pane", "")
key = d.get("key_sent", "")
rule = d.get("rule_name", "")
line_str = f" [{ts}] {act:<14} {ag.upper():<6} {sess:<12} ({pane}) -> sent '{key}' [{rule}]"
attr = self._attr("green") if act == "AUTO_APPROVED" else (self._attr("cyan") if "DRY" in act else self._attr("warn_banner"))
except Exception:
line_str = f" {l.strip()}"
attr = self._attr("normal")
self.safe_addstr(row_y, 2, line_str, attr)
# =================================================================
# Modals
# =================================================================
def _render_modals(self, h: int, w: int) -> None:
if not self.modal:
return
modal_w = min(68, w - 6)
modal_h = min(14, h - 4)
top_y = (h - modal_h) // 2
left_x = (w - modal_w) // 2
# Modal backdrop
for y in range(top_y, top_y + modal_h):
self.safe_addstr(y, left_x, " " * modal_w, self._attr("selected"))
# Border
self.safe_addstr(top_y, left_x, "┌" + "─" * (modal_w - 2) + "┐", self._attr("cyan"))
for y in range(top_y + 1, top_y + modal_h - 1):
self.safe_addstr(y, left_x, "│", self._attr("cyan"))
self.safe_addstr(y, left_x + modal_w - 1, "│", self._attr("cyan"))
self.safe_addstr(top_y + modal_h - 1, left_x, "└" + "─" * (modal_w - 2) + "┘", self._attr("cyan"))
if self.modal == "help":
self.safe_addstr(top_y + 1, left_x + 3, "󰋼 KEYBOARD CHEAT-SHEET", self._attr("header_sel"))
hints = [
"1-4 / F1-F4 / Tab: Switch tabs",
"j/k / Up/Down: Navigate rows",
"Space / a: Toggle Master Auto-Approvals",
"t: Toggle auto-approval for selected session",
"n: Spawn new tmux worker (Tab 2) / New Onboard (Tab 1)",
"o: Trigger one-shot approval check / Submit OTP",
"m: Test Regex Matcher with custom text",
"q / Esc: Exit modal or quit TUI",
]
for i, hint in enumerate(hints):
self.safe_addstr(top_y + 3 + i, left_x + 4, hint, self._attr("normal"))
elif self.modal == "spawn_worker":
self.safe_addstr(top_y + 1, left_x + 3, "SPAWN NEW TMUX WORKER", self._attr("header_sel"))
self.safe_addstr(top_y + 3, left_x + 3, "Enter session name & command:", self._attr("bold"))
self.safe_addstr(top_y + 5, left_x + 3, f"> {self.modal_input_buf}_", self._attr("cyan"))
self.safe_addstr(top_y + 7, left_x + 3, "Format: <session_name> [command] (e.g. dev-runner python3 worker.py)", self._attr("dim"))
self.safe_addstr(top_y + modal_h - 2, left_x + 3, "[Enter] Spawn [Esc] Cancel", self._attr("dim"))
elif self.modal == "test_regex":
self.safe_addstr(top_y + 1, left_x + 3, "EDIT TEST SAMPLE FOR REGEX MATCHER", self._attr("header_sel"))
self.safe_addstr(top_y + 3, left_x + 3, "Enter prompt excerpt to test:", self._attr("bold"))
self.safe_addstr(top_y + 5, left_x + 3, f"> {self.modal_input_buf[:55]}_", self._attr("cyan"))
self.safe_addstr(top_y + modal_h - 2, left_x + 3, "[Enter] Evaluate Match [Esc] Cancel", self._attr("dim"))
# =================================================================
# Input Handling
# =================================================================
def _handle_key(self, ch: int) -> bool:
if ch in (3, 4): # Ctrl+C or Ctrl+D
return False
if self.modal:
if ch in (27,): # Esc
self.modal = None
return True
if self.modal in ("spawn_worker", "test_regex"):
if ch in (curses.KEY_ENTER, 10, 13):
if self.modal == "spawn_worker":
parts = self.modal_input_buf.strip().split(maxsplit=1)
if parts:
sess = parts[0]
cmd = parts[1] if len(parts) > 1 else "bash"
run_tmux_cmd("/tmp/tmux-muse.sock", "new-session", "-d", "-s", sess, cmd)
self._refresh_tally()
self.set_toast(f"✔ Spawned worker '{sess}' running '{cmd}'", "success")
self.modal = None
elif self.modal == "test_regex":
self.test_text_buf = self.modal_input_buf
self._eval_test_match()
self.set_toast("Evaluated regex test text", "info")
self.modal = None
return True
elif ch in (curses.KEY_BACKSPACE, 127, 8):
self.modal_input_buf = self.modal_input_buf[:-1]
return True
elif 32 <= ch <= 126:
self.modal_input_buf += chr(ch)
return True
elif self.modal == "help":
self.modal = None
return True
return True
# Quit
if ch in (ord('q'), ord('Q')):
return False
# Help
if ch in (ord('?'), curses.KEY_F1):
self.modal = "help"
return True
# Tabs navigation: 1-4, F1-F4, Tab
if ch in (ord('1'),):
self.current_tab = 0
return True
elif ch in (ord('2'),):
self.current_tab = 1
return True
elif ch in (ord('3'),):
self.current_tab = 2
return True
elif ch in (ord('4'),):
self.current_tab = 3
return True
elif ch in (ord('\t'),): # Tab
self.current_tab = (self.current_tab + 1) % len(self.tabs)
return True
# Master Auto-Approve Toggle: Space or 'a'
if ch in (ord(' '), ord('a'), ord('A')) and self.current_tab in (1, 2):
self.approver_state.global_enabled = not self.approver_state.global_enabled
self.approver_state.save()
state_str = "ENABLED" if self.approver_state.global_enabled else "DISABLED"
self.set_toast(f"Master Auto-Approvals: {state_str}", "success" if self.approver_state.global_enabled else "warn")
self._refresh_tally()
return True
# Tab 0: Onboard Connects keys
if self.current_tab == 0:
if ch in (ord('j'), curses.KEY_DOWN):
self.sel_connect_idx = min(len(self.connects) - 1, self.sel_connect_idx + 1)
return True
elif ch in (ord('k'), curses.KEY_UP):
self.sel_connect_idx = max(0, self.sel_connect_idx - 1)
return True
elif ch in (ord('r'), ord('R')):
self._refresh_connects()
self.set_toast("Refreshed Onboard Connects", "info")
return True
elif ch in (ord('s'), ord('S')):
sel = self.connects[self.sel_connect_idx] if 0 <= self.sel_connect_idx < len(self.connects) else None
node = sel.get("node", "646") if sel else "646"
self.set_toast(f"Dispatched salvage work order for @{node}", "success")
return True
# Tab 1: Tmux Workers keys
elif self.current_tab == 1:
panes = self.tally.panes
if ch in (ord('j'), curses.KEY_DOWN):
self.sel_pane_idx = min(len(panes) - 1, self.sel_pane_idx + 1)
return True
elif ch in (ord('k'), curses.KEY_UP):
self.sel_pane_idx = max(0, self.sel_pane_idx - 1)
return True
elif ch in (ord('t'), ord('T')):
if panes and 0 <= self.sel_pane_idx < len(panes):
p = panes[self.sel_pane_idx]
cur = self.approver_state.sessions_enabled.get(p.session, True)
self.approver_state.sessions_enabled[p.session] = not cur
self.approver_state.save()
self._refresh_tally()
state_str = "ON" if not cur else "OFF"
self.set_toast(f"Toggled Auto-Approve for '{p.session}': {state_str}", "info")
return True
elif ch in (ord('n'), ord('N')):
self.modal = "spawn_worker"
self.modal_input_buf = ""
return True
elif ch in (ord('k'), ord('K')):
if panes and 0 <= self.sel_pane_idx < len(panes):
p = panes[self.sel_pane_idx]
run_tmux_cmd(p.socket, "kill-session", "-t", p.session)
self._refresh_tally()
self.set_toast(f"✔ Killed session '{p.session}'", "warn")
return True
elif ch in (ord('r'), ord('R')):
self._refresh_tally()
self.set_toast("Refreshed Tmux Workers", "info")
return True
# Tab 2: Rules & Auto-Approvals keys
elif self.current_tab == 2:
if ch in (ord('m'), ord('M')):
self.modal = "test_regex"
self.modal_input_buf = self.test_text_buf
return True
elif ch in (ord('o'), ord('O')):
res = self.runner.run_once()
self.set_toast(f"Executed single-pass check: {len(res)} action(s)", "success")
return True
elif ch in (ord('j'), curses.KEY_DOWN):
self.sel_rule_idx = min(len(self.approver_state.rules) - 1, self.sel_rule_idx + 1)
return True
elif ch in (ord('k'), curses.KEY_UP):
self.sel_rule_idx = max(0, self.sel_rule_idx - 1)
return True
# Tab 3: Logs keys
elif self.current_tab == 3:
if ch in (ord('j'), curses.KEY_DOWN):
self.sel_log_scroll = max(0, self.sel_log_scroll - 1)
return True
elif ch in (ord('k'), curses.KEY_UP):
self.sel_log_scroll += 1
return True
elif ch in (ord('e'), ord('E')):
self.set_toast("Synced status with https://box.muse-dev.online/ API", "success")
return True
return True
def _handle_mouse(self, mx: int, my: int, bstate: int) -> bool:
# Check tab clicks (my == 1)
if my == 1:
col = 1
for idx, tab_name in enumerate(self.tabs):
tab_w = len(tab_name) + 4
if col <= mx < col + tab_w:
self.current_tab = idx
return True
col += tab_w + 2
# Header Master Switch click (my == 0, right side)
h, w = self.stdscr.getmaxyx()
if my == 0 and mx >= w - 30:
self.approver_state.global_enabled = not self.approver_state.global_enabled
self.approver_state.save()
self._refresh_tally()
return True
return True
def run(self) -> None:
while True:
h, w = self.stdscr.getmaxyx()
self.stdscr.erase()
self._render_header(w)
if self.current_tab == 0:
self._render_tab_onboard(h, w)
elif self.current_tab == 1:
self._render_tab_tmux(h, w)
elif self.current_tab == 2:
self._render_tab_approvals(h, w)
else:
self._render_tab_logs(h, w)
self._render_footer(h, w)
self._render_modals(h, w)
self.stdscr.refresh()
try:
ch = self.stdscr.getch()
if ch != -1:
if ch == curses.KEY_MOUSE:
try:
_, mx, my, _, bstate = curses.getmouse()
self._handle_mouse(mx, my, bstate)
except Exception:
pass
else:
if not self._handle_key(ch):
break
except KeyboardInterrupt:
break
# Periodic background auto-approval check if enabled
now = time.time()
if self.approver_state.global_enabled and now - self.last_auto_poll > 2.0:
self.last_auto_poll = now
self.runner.run_once()
self._refresh_tally()
time.sleep(0.04)
def main() -> int:
try:
curses.wrapper(lambda stdscr: BoxOnboardTUI(stdscr).run())
finally:
try:
sys.stdout.write("\033[?1000l\033[?1002l\033[?1006l\033[?2004l")
sys.stdout.flush()
except Exception:
pass
return 0
if __name__ == "__main__":
sys.exit(main())
-74
View File
@@ -1,74 +0,0 @@
#!/usr/bin/env python3
"""
box-query.py — Lightweight client to query box.muse-dev.online APIs over HTTPS.
Works inside containers and nodes without direct SSH access:
Uses SSH signature authentication (?identity=bl&ts=...&sig=...) against
the VM Box API.
Usage:
box-query.py timers
box-query.py jobs
box-query.py agents
box-query.py dms [limit]
"""
import argparse
import json
import os
import subprocess
import sys
import time
import urllib.parse
import urllib.request
BOX_API_BASE = os.environ.get("BOX_API_BASE", "https://box.muse-dev.online")
BOX_SIGN_KEY = os.environ.get("BOX_SIGN_KEY", os.path.expanduser("~/.ssh/id_ed25519"))
def sign_request(identity, endpoint):
ts = str(int(time.time()))
payload = f"{ts}\n{endpoint}".encode()
try:
p = subprocess.run(
["ssh-keygen", "-Y", "sign", "-f", BOX_SIGN_KEY, "-n", "box"],
input=payload, capture_output=True, timeout=15)
if p.returncode != 0:
return None, None
return ts, p.stdout.decode()
except Exception:
return None, None
def query(endpoint):
ts, sig = sign_request("bl", endpoint)
if not ts or not sig:
sys.stderr.write("Failed to sign request (missing key or ssh-keygen error)\n")
sys.exit(1)
query_str = urllib.parse.urlencode({
"identity": "bl",
"ts": ts,
"sig": sig
})
url = f"{BOX_API_BASE}/api/box/{endpoint}?{query_str}"
req = urllib.request.Request(
url,
headers={"User-Agent": "NetVM-box-query/1.0 (container)"}
)
with urllib.request.urlopen(req, timeout=15) as resp:
data = resp.read().decode("utf-8")
try:
return json.loads(data)
except Exception:
return data
def main():
p = argparse.ArgumentParser(description="Query box.muse-dev.online APIs via HTTPS signature auth")
p.add_argument("endpoint", choices=["timers", "jobs", "agents", "dms", "health"])
p.add_argument("--json", action="store_true")
args = p.parse_args()
res = query(args.endpoint)
print(json.dumps(res, indent=2))
if __name__ == "__main__":
main()
+424
View File
@@ -0,0 +1,424 @@
#!/usr/bin/env python3
"""box-readback-loopback.py — Synchronous Readback Gate & Cognitive-Aware True Loopback Engine.
Architecture:
1. Readback Gate:
- Synchronously awaits agent confirmation (up to 45s) after task dispatch.
- Validates via Hybrid Tag ([READBACK] Ticket #...) + Semantic Fallback.
- Marks ticket 'in-progress' upon confirmation; marks 'blocked' and releases agent on timeout.
2. Cognitive-Aware True Loopbacks:
- Zero Token Burn: 100% silent while git commits or PRs are progressing.
- Cognitive Guard: Defer loopbacks while agent is THINKING / GENERATING.
- Dual-Layer Escalation:
* 15m inactive: Tier 1 non-intrusive Gitea ticket comment (@agent inquiry).
* 45m inactive: Tier 2 direct chat DM escalation (muse-cli-node send).
* 90m inactive: Tier 3 failure escalation (mark 'blocked', alert #lobby, release agent).
"""
import json
import os
import re
import socket
import subprocess
import sys
import time
import urllib.request
from datetime import datetime, timezone
from pathlib import Path
# Local imports
try:
import agent_cognitive_probe as acp
except ImportError:
acp = None
REMOTE_HOST = "100.123.153.75" # bl control node
DEFAULT_GITEA_URL = "https://tea.muse-dev.online"
LOOPBACK_STATE_FILE = Path("/tmp/box-loopback-state.json")
def is_running_on_bl():
try:
hn = socket.gethostname().lower()
if "bl" in hn:
return True
except Exception:
pass
return os.path.exists("/var/run/netns/warp-muse") or os.path.exists("/run/netns/warp-muse")
def get_gitea_token():
token = os.environ.get("GITEA_TOKEN", "3c26744525bceaf385aa09737f7e41af613627b6")
return token
def gitea_api(endpoint: str, method: str = "GET", data: dict = None):
token = get_gitea_token()
url = f"{DEFAULT_GITEA_URL}/api/v1{endpoint}"
headers = {
"Authorization": f"token {token}",
"Content-Type": "application/json",
"User-Agent": "Box-Work-CLI/1.0",
}
payload = json.dumps(data).encode("utf-8") if data else None
req = urllib.request.Request(url, data=payload, headers=headers, method=method)
try:
with urllib.request.urlopen(req, timeout=10) as r:
if r.status in (200, 201):
return json.loads(r.read().decode())
return {"status": r.status}
except urllib.error.HTTPError as e:
try:
return json.loads(e.read().decode())
except Exception:
return {"error": str(e), "code": e.code}
except Exception as e:
return {"error": str(e)}
# ----------------------------------------------------------------------
# CHAT COMMUNICATION HELPERS
# ----------------------------------------------------------------------
def send_agent_chat(agent: str, message: str) -> bool:
"""Delivers a message directly into the agent's web chat session."""
if is_running_on_bl():
cmd = ["/home/super/Projects/NetVM/bin/muse-cli-node", agent, "send", message]
else:
cmd = ["ssh", "-q", f"super@{REMOTE_HOST}",
f"/home/super/Projects/NetVM/bin/muse-cli-node {agent} send {subprocess.list2cmdline([message])}"]
try:
res = subprocess.run(cmd, capture_output=True, text=True, timeout=15)
return res.returncode == 0
except Exception:
return False
def get_agent_history(agent: str, limit: int = 5) -> list:
"""Retrieves recent chat messages from the agent's active session."""
if is_running_on_bl():
cmd = ["/home/super/Projects/NetVM/bin/muse-cli-node", agent, "history", "--limit", str(limit)]
else:
cmd = ["ssh", "-q", f"super@{REMOTE_HOST}",
f"/home/super/Projects/NetVM/bin/muse-cli-node {agent} history --limit {limit}"]
try:
res = subprocess.run(cmd, capture_output=True, text=True, timeout=15)
if res.returncode == 0 and res.stdout.strip():
data = json.loads(res.stdout.strip())
if isinstance(data, list):
return data
except Exception:
pass
return []
def get_latest_chat_seq(agent: str) -> int:
"""Finds the maximum sequence number in the agent's chat history."""
history = get_agent_history(agent, limit=3)
seqs = [m.get("seq", 0) for m in history if isinstance(m, dict) and "seq" in m]
return max(seqs) if seqs else 0
# ----------------------------------------------------------------------
# READBACK GATE
# ----------------------------------------------------------------------
def validate_readback(text: str, issue_num: int, agent: str) -> tuple[bool, str]:
"""Validates an agent readback using Hybrid Tag + Semantic Fallback.
Returns (is_valid, excerpt).
"""
if not text:
return False, ""
clean_text = text.strip()
issue_pattern = rf"#?{issue_num}\b"
# 1. Strict Tag Match: [READBACK] Ticket #<num> ...
if re.search(r"\[READBACK\]", clean_text, re.IGNORECASE) and re.search(issue_pattern, clean_text):
snippet = clean_text[:200].replace("\n", " ")
return True, snippet
# 2. Semantic Fallback: Mentions ticket number AND branch/accepted status
has_issue = bool(re.search(issue_pattern, clean_text))
has_branch_or_ack = bool(re.search(
rf"(dev/{agent}/|branch|accepted|working on|confirm|start(ed|ing)|received)",
clean_text, re.IGNORECASE
))
if has_issue and has_branch_or_ack:
snippet = clean_text[:200].replace("\n", " ")
return True, snippet
return False, ""
def wait_for_readback(agent: str, issue_num: int, initial_seq: int, timeout_s: int = 45, poll_s: float = 3.0) -> dict:
"""Synchronously polls for agent readback within timeout_s."""
start_time = time.time()
deadline = start_time + timeout_s
while time.time() < deadline:
elapsed = int(time.time() - start_time)
print(f"\r ⏳ Awaiting Readback from @{agent} ({elapsed}s / {timeout_s}s)...", end="", flush=True)
history = get_agent_history(agent, limit=4)
for msg in history:
seq = msg.get("seq", 0)
role = msg.get("role", "")
text = msg.get("text", "")
# Only check new assistant messages
if seq > initial_seq and role == "assistant":
valid, excerpt = validate_readback(text, issue_num, agent)
if valid:
print()
return {
"success": True,
"snippet": excerpt,
"elapsed": elapsed,
"seq": seq,
}
time.sleep(poll_s)
print()
return {
"success": False,
"timeout": True,
"elapsed": timeout_s,
}
def handle_readback_success(agent: str, issue_num: int, snippet: str):
"""Marks ticket in-progress and records confirmation on Gitea."""
# Label ticket in-progress
gitea_api(f"/repos/super/box/issues/{issue_num}/labels", method="POST", data={"labels": ["in-progress"]})
# Post confirmation comment
comment_body = f"🤖 **Readback Confirmed** by @{agent}:\n> {snippet}"
gitea_api(f"/repos/super/box/issues/{issue_num}/comments", method="POST", data={"body": comment_body})
def handle_readback_timeout(agent: str, issue_num: int, title: str):
"""Labels ticket blocked and unassigns agent so they return to IDLE."""
# Label ticket blocked
gitea_api(f"/repos/super/box/issues/{issue_num}/labels", method="POST", data={"labels": ["blocked"]})
# Post explanation comment
comment_body = (
f"⚠️ **Readback Timeout**: Agent @{agent} did not confirm ticket #{issue_num} "
f"within 45 seconds of dispatch. Releasing assignment to prevent deadlocks."
)
gitea_api(f"/repos/super/box/issues/{issue_num}/comments", method="POST", data={"body": comment_body})
# Unassign agent
gitea_api(f"/repos/super/box/issues/{issue_num}", method="PATCH", data={"assignees": []})
# ----------------------------------------------------------------------
# COGNITIVE-AWARE TRUE LOOPBACK ENGINE
# ----------------------------------------------------------------------
def load_loopback_state() -> dict:
if LOOPBACK_STATE_FILE.exists():
try:
with open(LOOPBACK_STATE_FILE) as f:
return json.load(f)
except Exception:
pass
return {}
def save_loopback_state(state: dict):
try:
with open(LOOPBACK_STATE_FILE, "w") as f:
json.dump(state, f, indent=2)
except Exception:
pass
def get_ticket_git_activity(agent: str, issue_num: int) -> datetime | None:
"""Checks the latest commit timestamp on the agent's branch dev/<agent>/<issue>-*."""
# Check Gitea branches for dev/<agent>/<issue_num>-*
branches = gitea_api("/repos/super/box/branches")
if isinstance(branches, list):
target_prefix = f"dev/{agent}/{issue_num}"
for b in branches:
name = b.get("name", "")
if target_prefix in name:
commit = b.get("commit", {})
ts_str = commit.get("timestamp")
if ts_str:
try:
return datetime.fromisoformat(ts_str.replace("Z", "+00:00"))
except Exception:
pass
return None
def get_ticket_last_activity(issue: dict, agent: str) -> tuple[datetime, str]:
"""Finds the most recent activity timestamp (git commit, comment, or issue creation)."""
issue_num = issue["number"]
latest_dt = datetime.fromisoformat(issue["created_at"].replace("Z", "+00:00"))
source = "issue_created"
# Check comments
comments = gitea_api(f"/repos/super/box/issues/{issue_num}/comments")
if isinstance(comments, list):
for c in comments:
c_dt = datetime.fromisoformat(c["created_at"].replace("Z", "+00:00"))
if c_dt > latest_dt:
latest_dt = c_dt
source = "gitea_comment"
# Check git branch commit
git_dt = get_ticket_git_activity(agent, issue_num)
if git_dt and git_dt > latest_dt:
latest_dt = git_dt
source = "git_commit"
return latest_dt, source
def run_loopback_sweep(dry_run: bool = False, verbose: bool = True) -> list:
"""Executes a single sweep of all open assigned tickets according to the 3-tier escalation model."""
now = datetime.now(timezone.utc)
state = load_loopback_state()
actions_taken = []
issues = gitea_api("/repos/super/box/issues?state=open")
if not isinstance(issues, list):
if verbose:
print("Failed to fetch open issues from Gitea.")
return []
assigned_issues = [i for i in issues if i.get("assignee")]
if verbose:
print(f"\n=== LOOPBACK SWEEP: {len(assigned_issues)} ACTIVE ASSIGNED TICKETS ({now.strftime('%H:%M:%SZ')}) ===")
for iss in assigned_issues:
issue_num = iss["number"]
title = iss.get("title", "")
agent = iss["assignee"]["username"]
key = str(issue_num)
ticket_state = state.get(key, {})
last_dt, source = get_ticket_last_activity(iss, agent)
inactive_s = (now - last_dt).total_seconds()
inactive_m = int(inactive_s // 60)
# Check cognitive state
cog = acp.get_passive_cognitive_state(agent) if acp else {"status": "IDLE", "cognitive_lock": False}
cog_status = cog.get("status", "IDLE")
is_thinking = cog_status == "THINKING" or cog.get("cognitive_lock")
if verbose:
print(f"Ticket #{issue_num} (@{agent}): {inactive_m}m inactive (source: {source}) | Cognitive: {cog_status}")
# Tier 0: Inactive < 15m or active git commits -> Complete silence
if inactive_m < 15 or source == "git_commit":
if verbose:
print(" 👉 Status: Active or within silent grace period (<15m). No action.")
continue
# Check cognitive guard: Defer if agent is thinking/generating
if is_thinking and cog_status != "INPUT_WAIT":
if verbose:
print(f" 🧠 Cognitive Guard: Deferring loopback — @{agent} is currently {cog_status}.")
continue
# Tier 1: 15m <= Inactivity < 45m -> Non-intrusive Gitea ticket comment
if 15 <= inactive_m < 45:
if ticket_state.get("tier1_sent"):
if verbose:
print(" 👉 Tier 1 comment already dispatched. Waiting for 45m threshold.")
continue
msg = (
f"🤖 @{agent} **Loopback Tier 1 Check** ({inactive_m}m elapsed):\n"
f"No git commits recorded on feature branch for Ticket #{issue_num}. "
f"Are you progressing or blocked? Reply with status or push a commit."
)
action_desc = f"Tier 1: Posted Gitea comment to #{issue_num} (@{agent})"
actions_taken.append(action_desc)
if not dry_run:
gitea_api(f"/repos/super/box/issues/{issue_num}/comments", method="POST", data={"body": msg})
ticket_state["tier1_sent"] = now.isoformat()
state[key] = ticket_state
save_loopback_state(state)
if verbose:
print(f" ✓ {action_desc}")
# Tier 2: 45m <= Inactivity < 90m -> Direct Chat DM Escalation
elif 45 <= inactive_m < 90:
if ticket_state.get("tier2_sent"):
if verbose:
print(" 👉 Tier 2 chat DM already dispatched. Waiting for 90m threshold.")
continue
chat_msg = (
f"[LOOPBACK ALERT] Ticket #{issue_num} ('{title}'): "
f"{inactive_m} minutes inactive with no git commits. "
f"Please confirm if blocked on tool execution, terminal approvals, or environment."
)
action_desc = f"Tier 2: Escalated to chat DM for @{agent} on #{issue_num}"
actions_taken.append(action_desc)
if not dry_run:
send_agent_chat(agent, chat_msg)
gitea_api(f"/repos/super/box/issues/{issue_num}/comments", method="POST", data={
"body": f"📣 **Loopback Tier 2 Escalation**: Inactivity reached {inactive_m}m. Sent direct chat DM to @{agent}."
})
ticket_state["tier2_sent"] = now.isoformat()
state[key] = ticket_state
save_loopback_state(state)
if verbose:
print(f" ✓ {action_desc}")
# Tier 3: Inactivity >= 90m (or unhandled INPUT_WAIT > 15m) -> Fail-closed escalation
elif inactive_m >= 90 or (cog_status == "INPUT_WAIT" and inactive_m >= 15):
action_desc = f"Tier 3: Ticket #{issue_num} marked BLOCKED; released @{agent} assignment"
actions_taken.append(action_desc)
if not dry_run:
# Label blocked
gitea_api(f"/repos/super/box/issues/{issue_num}/labels", method="POST", data={"labels": ["blocked"]})
# Post failure comment
gitea_api(f"/repos/super/box/issues/{issue_num}/comments", method="POST", data={
"body": (
f"🚨 **Loopback Tier 3 Escalation**: Inactivity reached {inactive_m}m with zero git progress. "
f"Ticket marked `blocked` and unassigned from @{agent} for operator intervention."
)
})
# Unassign agent
gitea_api(f"/repos/super/box/issues/{issue_num}", method="PATCH", data={"assignees": []})
ticket_state["tier3_sent"] = now.isoformat()
state[key] = ticket_state
save_loopback_state(state)
if verbose:
print(f" 🚨 {action_desc}")
if verbose:
print()
return actions_taken
def main():
import argparse
parser = argparse.ArgumentParser(description="Synchronous Readback & Cognitive True Loopback Engine")
sub = parser.add_subparsers(dest="cmd")
p_sweep = sub.add_parser("sweep", help="Run a loopback sweep across open tickets")
p_sweep.add_argument("--dry-run", action="store_true", help="Evaluate conditions without sending messages")
p_sweep.add_argument("--json", action="store_true", help="Output actions as JSON")
args = parser.parse_args()
if not args.cmd or args.cmd == "sweep":
dry_run = getattr(args, "dry_run", False)
as_json = getattr(args, "json", False)
actions = run_loopback_sweep(dry_run=dry_run, verbose=not as_json)
if as_json:
print(json.dumps({"ok": True, "actions": actions}, indent=2))
if __name__ == "__main__":
main()
-1072
View File
File diff suppressed because it is too large Load Diff
-1
View File
@@ -1 +0,0 @@
../watchers/box-stability-watcher.py
-204
View File
@@ -1,204 +0,0 @@
#!/usr/bin/env python3
"""
box-sys-op.py — Sandboxed execution helper for core system operations:
files.read, files.write, web.fetch, service.status, service.restart
Called via fixed argv from exec-constrained.py.
"""
import sys
import os
import json
import urllib.request
import urllib.error
import urllib.parse
import ipaddress
import subprocess
from pathlib import Path
REPO_ROOT = Path("/home/super/Projects/NetVM").resolve()
MAX_OUTPUT = 4096
ALLOWED_SERVICES = {
"board.service", "caddy.service", "response-harvester.timer",
"self-main-loop.timer", "job-heartbeat.timer", "job-scheduler.timer"
}
def safe_repo_path(raw):
clean = os.path.normpath(raw.strip())
if not os.path.isabs(clean):
clean = os.path.normpath(str(REPO_ROOT / clean))
real = Path(clean).resolve()
if not str(real).startswith(str(REPO_ROOT) + "/") and real != REPO_ROOT:
raise ValueError("path must reside inside repository root (/home/super/Projects/NetVM)")
return real
def op_files_read(path_str, max_lines=100):
p = safe_repo_path(path_str)
if not p.exists() or not p.is_file():
return {"ok": False, "error": f"File not found: {path_str}"}
with open(p, "r", encoding="utf-8", errors="replace") as f:
lines = f.readlines()
total_lines = len(lines)
snippet = "".join(lines[:max_lines])
truncated = total_lines > max_lines or len(snippet) > MAX_OUTPUT
if len(snippet) > MAX_OUTPUT:
snippet = snippet[:MAX_OUTPUT] + "\n... [truncated]"
return {
"ok": True,
"path": str(p.relative_to(REPO_ROOT)),
"lines": total_lines,
"displayed_lines": min(total_lines, max_lines),
"content": snippet,
"truncated": truncated
}
def op_files_write(path_str, content):
p = safe_repo_path(path_str)
p.parent.mkdir(parents=True, exist_ok=True)
with open(p, "w", encoding="utf-8") as f:
f.write(content)
return {
"ok": True,
"path": str(p.relative_to(REPO_ROOT)),
"bytes_written": len(content.encode("utf-8")),
"lines": content.count("\n") + 1
}
def op_web_fetch(url_str):
parsed = urllib.parse.urlparse(url_str)
if parsed.scheme not in ("http", "https"):
return {"ok": False, "error": "URL scheme must be http or https"}
host = parsed.hostname or ""
if not host or host in ("localhost", "127.0.0.1", "::1"):
return {"ok": False, "error": "Loopback destinations blocked"}
try:
ip = ipaddress.ip_address(host)
if ip.is_private or ip.is_loopback or ip.is_link_local:
return {"ok": False, "error": "Private and local IP addresses blocked"}
except ValueError:
pass
req = urllib.request.Request(
url_str,
headers={"User-Agent": "Mozilla/5.0 Box-Agent-Client/1.0"}
)
try:
with urllib.request.urlopen(req, timeout=10) as resp:
data = resp.read(MAX_OUTPUT + 1024).decode("utf-8", errors="replace")
status = resp.status
truncated = len(data) > MAX_OUTPUT
if truncated:
data = data[:MAX_OUTPUT] + "\n... [truncated]"
return {
"ok": True,
"url": url_str,
"status": status,
"length": len(data),
"body": data,
"truncated": truncated
}
except urllib.error.HTTPError as he:
return {"ok": False, "status": he.code, "error": f"HTTP {he.code}: {he.reason}"}
except Exception as e:
return {"ok": False, "error": str(e)}
def op_service_status(unit):
if unit not in ALLOWED_SERVICES:
return {"ok": False, "error": f"Service not allowed: {unit}"}
flag = "--user" if unit.endswith(".timer") else "--system"
cmd = ["systemctl", flag, "status", unit] if flag == "--user" else ["systemctl", "is-active", unit]
r = subprocess.run(cmd, capture_output=True, text=True, timeout=10)
active = "active" in r.stdout.lower() or "active" in r.stderr.lower()
return {
"ok": True,
"service": unit,
"active": active,
"status_line": r.stdout.splitlines()[0] if r.stdout.splitlines() else "unknown",
"output": r.stdout[:500].strip()
}
def op_service_restart(unit):
if unit not in ALLOWED_SERVICES:
return {"ok": False, "error": f"Service not allowed: {unit}"}
if unit.endswith(".timer"):
cmd = ["systemctl", "--user", "restart", unit]
else:
cmd = ["sudo", "-n", "systemctl", "restart", unit]
r = subprocess.run(cmd, capture_output=True, text=True, timeout=15)
return {
"ok": r.returncode == 0,
"service": unit,
"restarted": r.returncode == 0,
"error": r.stderr.strip() if r.returncode != 0 else None
}
def op_followup_schedule(args):
agent = args.get("agent")
if agent not in ("muse", "pip", "646", "opm", "dev", "def"):
return {"ok": False, "error": f"Invalid agent: {agent}"}
sender = args.get("sender") or agent
if sender not in ("muse", "pip", "646", "opm", "dev", "def"):
sender = agent
try:
in_m = float(args.get("in_m", 1))
except (TypeError, ValueError):
return {"ok": False, "error": "in_m must be a number"}
sec = max(5, int(in_m * 60))
prompt = args.get("prompt", "")
if not prompt or not isinstance(prompt, str):
return {"ok": False, "error": "prompt must be a non-empty string"}
if len(prompt) > 1000:
return {"ok": False, "error": "prompt exceeds 1000 characters"}
sidechat = args.get("thread") or args.get("sidechat")
cmd = [
"systemd-run", "--user", f"--on-active={sec}s",
sys.executable, str(REPO_ROOT / "bin" / "box-ctl.py"),
"notify", agent, prompt, "--sender", sender
]
if sidechat and isinstance(sidechat, str) and len(sidechat) <= 64:
cmd.extend(["--sidechat", sidechat])
r = subprocess.run(cmd, capture_output=True, text=True, timeout=15)
if r.returncode != 0:
return {"ok": False, "error": r.stderr.strip() or r.stdout.strip()}
return {
"ok": True,
"agent": agent,
"in_seconds": sec,
"sidechat": sidechat,
"timer_info": (r.stderr or r.stdout).strip()
}
def main():
if len(sys.argv) < 2:
print(json.dumps({"ok": False, "error": "missing operation"}))
sys.exit(1)
op = sys.argv[1]
raw_args = sys.stdin.read()
try:
args = json.loads(raw_args) if raw_args.strip() else {}
except Exception as e:
print(json.dumps({"ok": False, "error": f"bad json args: {e}"}))
sys.exit(1)
try:
if op == "files.read":
res = op_files_read(args.get("path", ""), int(args.get("lines", 100)))
elif op == "files.write":
res = op_files_write(args.get("path", ""), args.get("content", ""))
elif op == "web.fetch":
res = op_web_fetch(args.get("url", ""))
elif op == "service.status":
res = op_service_status(args.get("unit", args.get("name", "")))
elif op == "service.restart":
res = op_service_restart(args.get("unit", args.get("name", "")))
elif op in ("followup.create", "followup.schedule"):
res = op_followup_schedule(args)
else:
res = {"ok": False, "error": f"unknown operation: {op}"}
except Exception as e:
res = {"ok": False, "error": str(e)}
print(json.dumps(res))
if __name__ == "__main__":
main()
-1
View File
@@ -1 +0,0 @@
muse-tui.py
+1138
View File
File diff suppressed because it is too large Load Diff
+1
View File
@@ -0,0 +1 @@
box-readback-loopback.py
+1
View File
@@ -0,0 +1 @@
/home/super/Projects/NetVM/bin/box-work.py
-413
View File
@@ -1,413 +0,0 @@
"""
brain.py — Main-loop brain workspace module.
The brain sidechat ("main-loop brain" on opm's account) is where the main loop
OPERATES: it posts its thinking there, reads operator instructions from there,
and keeps its working state visible there.
Deploy to: ~/Projects/NetVM/bin/brain.py on bl (alongside self_main_loop.py).
Design source: ~/workspace/main-loop-brain-design.md
User directive 2026-10-04: "we need main loop to operate its brains in side chat"
SAFETY CONTRACT (do not weaken):
- Brain posts carry a [BRAIN <ts>] marker, NEVER a [JOB <id>] marker.
The response-harvester keys off [JOB ...]; a thinking note must never
look actionable.
- Brain posts NEVER use dm.py --expect-reply. Thinking creates no followup
records, no nudges, no escalations.
- When reading the brain, the loop skips its own messages (sender check).
The loop must never digest its own thinking as agent activity.
- !loop commands are honored ONLY from AUTHORIZED_SENDERS. Everything else
is read as context, never as instruction.
- Malformed commands get a one-line correction posted to the brain.
Never silent, never a crash.
- Cap: MAX_BRAIN_POSTS_PER_TICK posts per tick. The brain must not amplify.
State lives in the existing watermark JSON file under the "brain" key:
{"brain": {"ignores": {"646": <expires_epoch>}, "quiet_until": <epoch|0>,
"brain_watermark": <epoch>, "tick": <int>}}
Absolute expiries so a dead loop cannot leave an agent ignored forever.
"""
import json
import os
import re
import subprocess
import time
# ---------------------------------------------------------------------------
# Constants
# ---------------------------------------------------------------------------
BRAIN_AGENT = "opm"
BRAIN_SIDECHAT_NAME = "main-loop brain"
# Operator identities allowed to issue !loop commands. These must match the
# DM-signer / board identity strings; do not invent new ones here.
AUTHORIZED_SENDERS = {"super", "operator-646", "operator-main"}
# Sender identities the loop itself posts under (skipped on read-back).
OWN_SENDERS = {"main-loop", "self_main_loop", "operator-main-loop", "opm"}
BRAIN_MARKER_RE = re.compile(r"\[BRAIN\s+([^\]]+)\]")
JOB_MARKER_RE = re.compile(r"\[JOB\s+([^\]]+)\]")
LOOP_CMD_RE = re.compile(r"^\s*!loop\s+(\S+)(.*)$", re.IGNORECASE)
MAX_BRAIN_POSTS_PER_TICK = 4
DEFAULT_BIN_DIR = os.path.expanduser("~/Projects/NetVM/bin")
VALID_AGENTS = {"muse", "pip", "646", "opm"}
DUR_RE = re.compile(r"^\s*(\d+)\s*([smh])\s*$", re.IGNORECASE)
# ---------------------------------------------------------------------------
# Pure logic — fully unit-testable, no I/O
# ---------------------------------------------------------------------------
def parse_duration(text):
"""'30m' -> 1800.0, '1h' -> 3600.0, '90s' -> 90.0. None if malformed."""
m = DUR_RE.match(text or "")
if not m:
return None
n, unit = int(m.group(1)), m.group(2).lower()
return float(n * {"s": 1, "m": 60, "h": 3600}[unit])
def is_own_message(sender):
s = (sender or "").strip().lower()
return s in {x.lower() for x in OWN_SENDERS}
def is_authorized(sender):
return (sender or "").strip() in AUTHORIZED_SENDERS
def extract_loop_commands(messages):
"""messages: list of {"sender": str, "text": str, "ts": float}.
Returns [(sender, verb, args, msg)] for !loop lines from any sender
(authorization is applied by the caller so corrections can name names)."""
cmds = []
for msg in messages:
text = msg.get("text") or ""
for line in text.splitlines():
m = LOOP_CMD_RE.match(line)
if m:
cmds.append((msg.get("sender", "?"),
m.group(1).lower(),
m.group(2).strip(),
msg))
return cmds
def prune_expired(state, now=None):
"""Drop expired ignores / quiet. Returns (notes, changed)."""
now = now if now is not None else time.time()
notes, changed = [], False
ignores = state.setdefault("ignores", {})
for agent in list(ignores):
if ignores[agent] <= now:
del ignores[agent]
notes.append("resuming %s (ignore expired)" % agent)
changed = True
if state.get("quiet_until", 0) and state["quiet_until"] <= now:
state["quiet_until"] = 0
notes.append("quiet period ended, prompts resumed")
changed = True
return notes, changed
def is_ignored(state, agent, now=None):
now = now if now is not None else time.time()
return state.get("ignores", {}).get(agent, 0) > now
def is_quiet(state, now=None):
now = now if now is not None else time.time()
return (state.get("quiet_until", 0) or 0) > now
def apply_command(sender, verb, args, state, now=None):
"""Apply one !loop command. Returns (ack_text, changed)."""
now = now if now is not None else time.time()
changed = False
if verb == "ignore":
parts = args.split()
if len(parts) != 2 or parts[0] not in VALID_AGENTS:
return ("usage: !loop ignore <agent> <dur> (agent: %s, dur like 30m/1h)"
% "/".join(sorted(VALID_AGENTS)), False)
dur = parse_duration(parts[1])
if dur is None or dur <= 0:
return ("bad duration %r, try 30m or 1h" % parts[1], False)
state.setdefault("ignores", {})[parts[0]] = now + dur
return ("ignoring %s for %s (until %s)"
% (parts[0], parts[1],
time.strftime("%H:%M UTC", time.gmtime(now + dur))), True)
if verb == "unignore":
agent = args.strip()
if agent not in VALID_AGENTS:
return ("usage: !loop unignore <agent>", False)
if state.get("ignores", {}).pop(agent, None) is not None:
changed = True
return ("resuming %s now" % agent, True)
return ("%s was not ignored" % agent, False)
if verb == "quiet":
dur = parse_duration(args)
if dur is None or dur <= 0:
return ("usage: !loop quiet <dur> (dur like 30m/1h)", False)
state["quiet_until"] = now + dur
changed = True
return ("quiet for %s: reads continue, digests land here, "
"no per-agent escalation" % args.strip(), True)
if verb == "unquiet":
if state.get("quiet_until"):
state["quiet_until"] = 0
changed = True
return ("prompts resumed", True)
return ("was not quiet", False)
if verb == "check":
# The tick loop honors this by running the agent reads immediately
# rather than waiting for the next timer fire. Caller sets the flag.
return ("CHECK_REQUESTED", True)
if verb == "status":
return ("STATUS_REQUESTED", False)
return ("unknown command %r, try: ignore, unignore, quiet, unquiet, "
"check, status" % verb, False)
def format_thinking(tick, summary, state, now=None):
"""summary: dict with keys seen{agent:(new,q,urgent)}, escalated[JOB ids],
skipped_info[int], errors[int], closure_rate[float|None]."""
now = now if now is not None else time.time()
ts = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime(now))
seen_bits = []
for agent in ("646", "pip", "muse", "opm"):
new, q, u = summary.get("seen", {}).get(agent, (0, 0, 0))
bit = "%s:%dnew" % (agent, new)
if q:
bit += "(%d?)" % q
if u:
bit += "(%d!)" % u
seen_bits.append(bit)
esc = summary.get("escalated", [])
lines = [
"[BRAIN %s tick=%d]" % (ts, tick),
"seen: " + ", ".join(seen_bits),
"decided: escalated %d%s | skipped %d info | ignored: %s" % (
len(esc),
(" (%s)" % ", ".join(esc[:3])) if esc else "",
summary.get("skipped_info", 0),
", ".join(sorted(state.get("ignores", {}))) or "none"),
"errors: %d | quiet: %s | closure: %s" % (
summary.get("errors", 0),
"yes" if is_quiet(state, now) else "no",
("%.2f" % summary["closure_rate"])
if summary.get("closure_rate") is not None else "n/a"),
]
return "\n".join(lines)
def format_status(tick, cfg, state, watermarks, health, now=None):
now = now if now is not None else time.time()
ts = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime(now))
lines = ["[BRAIN %s tick=%d] status" % (ts, tick)]
lines.append("agents: " + ", ".join(
"%s=%s" % (a, "on" if cfg.get(a) else "off")
for a in ("muse", "pip", "646", "opm")))
ign = state.get("ignores", {})
lines.append("ignores: " + (", ".join(
"%s until %s" % (a, time.strftime("%H:%M UTC", time.gmtime(e)))
for a, e in sorted(ign.items())) or "none"))
q = state.get("quiet_until", 0)
lines.append("quiet: " + ("until %s" % time.strftime("%H:%M UTC", time.gmtime(q))
if q and q > now else "no"))
ages = []
for a in ("muse", "pip", "646", "opm"):
wm = (watermarks or {}).get(a, 0)
ages.append("%s:%dm" % (a, int((now - wm) / 60)) if wm else "%s:never" % a)
lines.append("watermark age: " + ", ".join(ages))
if health:
lines.append("digest health: delivered=%d acked=%d closed=%d stale=%d "
"closure=%.2f" % (
health.get("delivered", 0), health.get("acked", 0),
health.get("closed", 0), health.get("stale", 0),
health.get("closure_rate", 0.0)))
return "\n".join(lines)
# ---------------------------------------------------------------------------
# I/O adapters — thin shells over the bl tools. Verify paths on bl.
# ---------------------------------------------------------------------------
class BrainIO:
"""Send/receive against the brain sidechat.
send: dm.py, WITHOUT --expect-reply (thinking is never actionable).
read: pluggable read_fn(messages-since-ts); defaults to None and must be
wired to the same chat-read primitive the response-harvester uses.
"""
def __init__(self, bin_dir=DEFAULT_BIN_DIR, read_fn=None,
sidechat_map_path=None):
self.bin_dir = bin_dir
self.read_fn = read_fn
self.sidechat_map_path = (sidechat_map_path or
os.path.join(bin_dir, "..",
"job-sidechats.json"))
self._posts_this_tick = 0
# -- name resolution: never hardcode the thread UUID -------------------
def brain_thread_uuid(self):
"""Resolve 'main-loop brain' -> UUID via job-sidechats.json."""
try:
with open(os.path.normpath(self.sidechat_map_path)) as f:
data = json.load(f)
except (OSError, ValueError):
return None
# schema: {"sidechats": {"main-loop brain": {"uuid": ...}}} or flat
node = data.get("sidechats", data).get(BRAIN_SIDECHAT_NAME)
if isinstance(node, dict):
return node.get("uuid") or node.get("thread_uuid")
return node if isinstance(node, str) else None
# -- write --------------------------------------------------------------
def reset_tick_budget(self):
self._posts_this_tick = 0
def post(self, text):
"""Post thinking/acks to the brain. Returns True on VERIFIED send."""
if self._posts_this_tick >= MAX_BRAIN_POSTS_PER_TICK:
return False
if JOB_MARKER_RE.search(text):
raise ValueError("refusing to post [JOB ...] to the brain")
dm = os.path.join(self.bin_dir, "dm.py")
# Same path the loop uses for opm digests; NO --expect-reply:
# thinking is never actionable and must not create followups.
cmd = [dm, "send", "--agent", BRAIN_AGENT,
"--to", BRAIN_AGENT, "--target", BRAIN_SIDECHAT_NAME,
"--message", text]
try:
p = subprocess.run(cmd, capture_output=True, text=True,
timeout=120)
except (OSError, subprocess.TimeoutExpired):
return False
ok = "VERIFIED" in (p.stdout or "")
if ok:
self._posts_this_tick += 1
return ok
# -- read ---------------------------------------------------------------
def read_new(self, since_ts):
"""Return [{"sender","text","ts"}] newer than since_ts. Skips own."""
if self.read_fn is None:
return []
try:
msgs = self.read_fn(BRAIN_AGENT, BRAIN_SIDECHAT_NAME, since_ts) or []
except Exception:
return []
return [m for m in msgs
if (m.get("ts", 0) or 0) > since_ts
and not is_own_message(m.get("sender"))]
# ---------------------------------------------------------------------------
# Workspace — one object per tick
# ---------------------------------------------------------------------------
class BrainWorkspace:
"""Owns the brain side of a main-loop tick.
Usage in self_main_loop.py tick():
brain = BrainWorkspace(watermark_path, read_fn=<harvester reader>)
cmds_outcome = brain.intake() # read, parse, apply, ack
... existing per-agent reads, skipping brain.ignored(agent) ...
brain.post_thinking(summary) # the loop's reasoning, visible
brain.save()
"""
def __init__(self, watermark_path, read_fn=None, bin_dir=DEFAULT_BIN_DIR):
self.watermark_path = watermark_path
self.io = BrainIO(bin_dir=bin_dir, read_fn=read_fn)
self._data = self._load()
self.state = self._data.setdefault("brain", {})
self.io.reset_tick_budget()
self.check_requested = False
self.status_requested = False
# -- persistence ---------------------------------------------------------
def _load(self):
try:
with open(self.watermark_path) as f:
return json.load(f)
except (OSError, ValueError):
return {}
def save(self):
tmp = self.watermark_path + ".tmp"
with open(tmp, "w") as f:
json.dump(self._data, f, indent=2)
os.replace(tmp, self.watermark_path)
# -- tick intake: read -> parse -> apply -> ack --------------------------
def intake(self):
"""Process new brain messages. Returns dict of what happened."""
now = time.time()
outcome = {"commands": 0, "acks": 0, "expired_notes": 0,
"ignored_senders": 0}
since = float(self.state.get("brain_watermark", 0))
msgs = self.io.read_new(since)
newest = since
for m in msgs:
newest = max(newest, float(m.get("ts", 0) or 0))
if msgs:
self.state["brain_watermark"] = newest
notes, _ = prune_expired(self.state, now)
for n in notes:
if self.io.post("[BRAIN %s] %s" % (
time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime(now)), n)):
outcome["expired_notes"] += 1
for sender, verb, args, _msg in extract_loop_commands(msgs):
if not is_authorized(sender):
outcome["ignored_senders"] += 1
continue
outcome["commands"] += 1
ack, _changed = apply_command(sender, verb, args, self.state, now)
if ack == "CHECK_REQUESTED":
self.check_requested = True
ack = "out-of-cycle check armed for this tick"
elif ack == "STATUS_REQUESTED":
self.status_requested = True
continue # status posts at end of tick with full context
if self.io.post("[BRAIN %s] @%s %s" % (
time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime(now)),
sender, ack)):
outcome["acks"] += 1
return outcome
# -- state queries for the tick loop --------------------------------------
def ignored(self, agent):
return is_ignored(self.state, agent)
def quiet(self):
return is_quiet(self.state)
# -- end of tick -----------------------------------------------------------
def post_thinking(self, summary, cfg=None, watermarks=None, health=None):
tick = int(self.state.get("tick", 0)) + 1
self.state["tick"] = tick
ok = self.io.post(format_thinking(tick, summary, self.state))
if self.status_requested and cfg is not None:
self.io.post(format_status(tick, cfg, self.state,
watermarks, health))
self.status_requested = False
return ok
-89
View File
@@ -1,89 +0,0 @@
#!/usr/bin/env python3
"""
Bridge CLI: Front-Door ↔ muse.ai metadata bridge
Manages mappings between front-door channels and muse.ai side chats.
Storage: ~/Projects/NetVM/bridge/mappings.json (operator-managed via SSH)
Usage:
bridge.py list --channel #ops
bridge.py add --agent 646 --chat <id> --channel #ops --purpose "incident-123"
bridge.py remove <bridge_id>
"""
import json
import argparse
from pathlib import Path
from datetime import datetime, timezone
BRIDGE_DIR = Path.home() / "Projects" / "NetVM" / "bridge"
MAPPINGS_FILE = BRIDGE_DIR / "mappings.json"
def load_mappings():
if not MAPPINGS_FILE.exists():
return []
with open(MAPPINGS_FILE) as f:
return json.load(f)
def save_mappings(mappings):
BRIDGE_DIR.mkdir(parents=True, exist_ok=True)
with open(MAPPINGS_FILE, 'w') as f:
json.dump(mappings, f, indent=2)
def cmd_list(args):
mappings = load_mappings()
if args.channel:
mappings = [m for m in mappings if m['frontdoor_channel'] == args.channel]
if not mappings:
print("No mappings found.")
return
for m in mappings:
print(f"{m['bridge_id']}: {m['agent']}/{m['muse_side_chat_id'][:8]} -> {m['frontdoor_channel']} ({m['purpose']}) [{m['status']}]")
def cmd_add(args):
mappings = load_mappings()
bridge_id = f"br-{datetime.now(timezone.utc).strftime('%Y%m%d%H%M%S')}"
mapping = {
"bridge_id": bridge_id,
"frontdoor_channel": args.channel,
"muse_side_chat_id": args.chat,
"agent": args.agent,
"linked_at": datetime.now(timezone.utc).isoformat(),
"linked_by": args.by or "operator",
"purpose": args.purpose,
"status": "active"
}
mappings.append(mapping)
save_mappings(mappings)
print(f"Added {bridge_id}")
def cmd_remove(args):
mappings = load_mappings()
mappings = [m for m in mappings if m['bridge_id'] != args.bridge_id]
save_mappings(mappings)
print(f"Removed {args.bridge_id}")
def main():
p = argparse.ArgumentParser(description="Bridge: front-door to muse.ai")
sub = p.add_subparsers(dest='cmd', required=True)
p_list = sub.add_parser('list', help='List mappings')
p_list.add_argument('--channel', help='Filter by front-door channel')
p_list.set_defaults(func=cmd_list)
p_add = sub.add_parser('add', help='Add a mapping')
p_add.add_argument('--agent', required=True, help='Agent (muse/pip/646)')
p_add.add_argument('--chat', required=True, help='muse.ai side chat ID')
p_add.add_argument('--channel', required=True, help='Front-door channel (#ops, #lobby, etc.)')
p_add.add_argument('--purpose', required=True, help='Purpose of the side chat')
p_add.add_argument('--by', help='Who linked (operator/agent)')
p_add.set_defaults(func=cmd_add)
p_rm = sub.add_parser('remove', help='Remove a mapping')
p_rm.add_argument('bridge_id', help='Bridge ID to remove')
p_rm.set_defaults(func=cmd_remove)
args = p.parse_args()
args.func(args)
if __name__ == '__main__':
main()
-51
View File
@@ -1,51 +0,0 @@
#!/bin/bash
# cdp-latency-check.sh — measure CDP relay latency per node.
#
# Probes each node's host-side relay (PEER_IP:CDP_PORT from netvm-names.sh,
# which pins the registry ports) at /json/version and prints one line per
# node:
# name:latency_ms:code
# On failure (timeout / connection refused / non-200-ish transport error):
# name:FAIL:000
#
# Self-contained: only needs netvm-names.sh in the same directory and curl.
set -u
BIN_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# shellcheck disable=SC1091
source "$BIN_DIR/netvm-names.sh"
TIMEOUT_S=10
# Registry-driven node list (was hardcoded 4 nodes; def/dev had no
# cdp-latency coverage — 2026-10-06).
watched_nodes() {
"$BIN_DIR/netvm-registry.py" 2>/dev/null | cut -d: -f1
}
# Allow sourcing for tests without running checks.
if [ "${CDP_LATENCY_CHECK_LIB_ONLY:-}" = "1" ]; then
return 0 2>/dev/null || exit 0
fi
NODES="$(watched_nodes)"
if [ -z "$NODES" ]; then
echo "node registry empty/unreadable" >&2
exit 1
fi
# shellcheck disable=SC2086 (intended word splitting: one node per word)
for node in $NODES; do
netvm_names "$node"
url="http://${PEER_IP}:${CDP_PORT}/json/version"
probe=$(curl -s -m "$TIMEOUT_S" -o /dev/null -w "%{time_total} %{http_code}" "$url" 2>/dev/null)
rc=$?
secs=$(printf '%s' "$probe" | awk '{print $1}')
code=$(printf '%s' "$probe" | awk '{print $2}')
if [ "$rc" -ne 0 ] || [ -z "$code" ] || [ "$code" = "000" ]; then
echo "${node}:FAIL:000"
continue
fi
ms=$(awk -v s="$secs" 'BEGIN { printf "%d", (s + 0) * 1000 }')
echo "${node}:${ms}:${code}"
done
-145
View File
@@ -1,145 +0,0 @@
#!/usr/bin/env bash
# cdp-relay-watchdog.sh — keep per-node CDP relays alive and correctly routed.
# Two-stage health check:
# 1. Host veth IP must be assigned (veth_healthy). Without it the relay is
# unreachable from the host no matter how many times we restart it —
# observed 2026-10-04 (muse/pip veths existed but had no IPs). FAIL_LOUD
# in the log; do NOT auto-fix (veth recreation touches WireGuard/iptables).
# 2. Relay connectivity (relay_healthy): curl to veth IP:port, not pidfile
# (which goes stale and lies — observed 2026-10-04).
# If a relay is down or misrouted: kill it and restart via the exact
# netvm-node-up.sh relay invocation inside the node's netns.
# Runs every 5 min via systemd timer cdp-relay-watchdog.timer.
# Pattern mirrors chromebox-watchdog.sh (stage-specific logging, rotation).
set -euo pipefail
LOCK="/tmp/cdp-relay-watchdog.lock"
# Tests source this file with CDP_RELAY_WATCHDOG_LIB_ONLY=1: they call
# helpers without running checks, so no lock is needed.
if [ "${CDP_RELAY_WATCHDOG_LIB_ONLY:-}" != "1" ]; then
exec 9>"$LOCK"
if ! flock -n 9; then
echo "[$(date -u +%FT%TZ)] another relay watchdog run in progress, skipping" >&2
exit 0
fi
fi
NETVM_BIN="/home/super/Projects/NetVM/bin"
LOG="/home/super/Projects/NetVM/cdp-relay-watchdog.log"
rotate_log() {
local f="$1"
[ -f "$f" ] || return 0
local sz
sz=$(stat -c%s "$f" 2>/dev/null || echo 0)
if [ "$sz" -gt 10485760 ]; then
mv -f "$f" "$f.1"
echo "[$(date -u +%FT%TZ)] log rotated" > "$f"
fi
}
rotate_log "$LOG"
log() { echo "[$(date -u +%FT%TZ)] $*" | tee -a "$LOG"; }
# node -> "veth_ip:port" via netvm-names.sh (hash-derived, don't hardcode)
relay_target() {
local node="$1" reg_port=""
# shellcheck disable=SC1091
. "$NETVM_BIN/netvm-names.sh"
netvm_names "$node" || return 1
# The registry is the source of truth for ports (new nodes propagate
# automatically); netvm-names pinning is the fallback.
if reg_port=$("$NETVM_BIN/netvm-registry.py" "$node" 2>/dev/null); then
[ -n "$reg_port" ] && CDP_PORT="$reg_port"
fi
echo "$PEER_IP:$CDP_PORT"
}
node_port() { echo "${1##*:}"; }
# node -> "VETH GW" via netvm-names.sh
node_veth() {
local node="$1"
# shellcheck disable=SC1091
. "$NETVM_BIN/netvm-names.sh"
netvm_names "$node" || return 1
echo "$VETH $GW"
}
# Is the host-side veth IP assigned? The relay listens on the netns-side peer
# IP; the host reaches it via the veth interface's GW address. If the GW IP is
# missing, the relay is unreachable from the host — restarting the relay is
# pointless and masks the real problem.
veth_healthy() {
local node="$1" veth gw
read -r veth gw <<< "$(node_veth "$node")" || return 1
ip addr show dev "$veth" 2>/dev/null | grep -q "inet ${gw}/" || return 1
return 0
}
relay_healthy() {
local node="$1" target
target="$(relay_target "$node")" || return 1
curl -s -m 8 "http://$target/json/version" 2>/dev/null | grep -q '"Browser"' || return 1
return 0
}
restart_relay() {
local node="$1" target port veth_ip netns
target="$(relay_target "$node")"
veth_ip="${target%%:*}"
port="${target##*:}"
netns="warp-$node"
# Kill any existing relay for this node's port (correct or not)
# Relays are root-owned (started via sudo ip netns exec); the timer runs as
# super, so the kill needs sudo too. Without it pkill fails EPERM silently
# and the "restart" false-positives via SO_REUSEADDR double-bind.
sudo -n pkill -f "netvm-cdp-relay.py .* $port 127.0.0.1 $port" 2>/dev/null || true
sleep 2
# Launch inside the netns, listening on the veth IP (host-reachable)
sudo -n ip netns exec "$netns" setsid nohup python3 \
"$NETVM_BIN/netvm-cdp-relay.py" "$veth_ip" "$port" 127.0.0.1 "$port" \
>>"$LOG" 2>&1 < /dev/null &
sleep 5
if relay_healthy "$node"; then
log "[$node] relay restarted OK on $target"
return 0
else
log "[$node] relay restart FAILED on $target — needs operator attention"
return 1
fi
}
# Registry-driven node list: every active node gets relay supervision
# (the old hardcoded 4-node list left def/dev unsupervised — 2026-10-06).
watched_nodes() {
"$NETVM_BIN/netvm-registry.py" 2>/dev/null | cut -d: -f1
}
# Allow sourcing for tests without running checks.
if [ "${CDP_RELAY_WATCHDOG_LIB_ONLY:-}" = "1" ]; then
return 0 2>/dev/null || exit 0
fi
FAILED=0
NODES="$(watched_nodes)"
if [ -z "$NODES" ]; then
log "FAIL_LOUD: node registry empty/unreadable, skipping run"
exit 1
fi
# shellcheck disable=SC2086 (intended word splitting: one node per word)
for node in $NODES; do
# Stage 1: host veth IP. Fail loud, skip relay restart (pointless).
if ! veth_healthy "$node"; then
read -r veth gw <<< "$(node_veth "$node")"
log "[$node] FAIL_LOUD: host veth $veth missing IP $gw — relay unreachable, needs netvm-node-up.sh $node (manual)"
FAILED=1
continue
fi
# Stage 2: relay connectivity.
if relay_healthy "$node"; then
continue
fi
target="$(relay_target "$node")"
log "[$node] relay unhealthy on $target, restarting"
restart_relay "$node" || FAILED=1
done
exit $FAILED
-239
View File
@@ -1,239 +0,0 @@
#!/usr/bin/env python3
"""
Per-browser CDP operation queue for the NetVM fleet.
Problem: nothing coordinates browser operations. DM sends (dm.py), tab
operations, agent reads (muse-chat-api.py), and watchdog restarts all hit
the same Chromium with zero scheduling. Functional tests pass under load,
but there is no backpressure — under real fleet concurrency this will
overwhelm the machine.
Solution: per-node FIFO queue with priority levels and a cap on concurrent
CDP operations per browser.
Cross-process design: dm.py shells out to muse-chat-api.py via subprocess,
so in-process locks (threading.Semaphore) alone cannot coordinate. This
module uses flock'd slot files + ticket files in /tmp, which work across
processes AND threads.
Usage:
from cdp_queue import cdp_slot, PRIORITY_HIGH
with cdp_slot("opm", priority=PRIORITY_HIGH):
... do CDP work ...
Explicit acquire/release:
from cdp_queue import acquire, release, QueueTimeout
token = acquire("opm", priority=PRIORITY_HIGH, timeout=60)
try:
...
finally:
release(token)
Priority levels (lower number = higher priority):
PRIORITY_HIGH = 0 # DM sends (user-facing)
PRIORITY_NORMAL = 1 # tab opens, reads
PRIORITY_LOW = 2 # background scans
Fairness: tickets are ordered by (priority, arrival time). A waiter only
proceeds when its ticket is first in line AND a slot is free. No starvation:
a low-priority ticket eventually becomes the oldest and gets served.
"""
import fcntl
import logging
import os
import time
import uuid
log = logging.getLogger("cdp_queue")
# ---- Tunables ----
PRIORITY_HIGH = 0
PRIORITY_NORMAL = 1
PRIORITY_LOW = 2
MAX_CONCURRENT = 2 # max simultaneous CDP ops per browser
ACQUIRE_TIMEOUT = 60.0 # fail loud instead of hanging forever
WARN_AFTER = 10.0 # log a warning when a waiter waits this long
POLL_INTERVAL = 0.05 # ticket/slot poll cadence
QUEUE_DIR = "/tmp/cdp-queue"
VALID_NODES = ("muse", "pip", "646", "opm", "dev", "def")
class QueueTimeout(Exception):
"""Raised when a slot cannot be acquired within the timeout."""
pass
def _node_dir(node):
return os.path.join(QUEUE_DIR, node)
def _tickets_dir(node):
return os.path.join(_node_dir(node), "tickets")
def _slot_path(node, i):
return os.path.join(_node_dir(node), "slot-%d.lock" % i)
def _ensure_dirs(node):
os.makedirs(_tickets_dir(node), exist_ok=True)
# Pre-create slot files so flock targets always exist
for i in range(MAX_CONCURRENT):
p = _slot_path(node, i)
if not os.path.exists(p):
open(p, "a").close()
def _read_tickets(node):
"""Return sorted list of (priority, timestamp, ticket_name), purging stale tickets."""
tdir = _tickets_dir(node)
out = []
now = time.time()
try:
for name in os.listdir(tdir):
if not name.endswith(".ticket"):
continue
try:
# ticket name: "<prio>-<timestamp>-<uuid>.ticket"
parts = name[:-7].split("-")
prio_s, ts_s = parts[0], parts[1]
ts = float(ts_s)
# Purge stale ticket if process crashed or timed out ungracefully
if now - ts > (ACQUIRE_TIMEOUT * 2):
try:
os.unlink(os.path.join(tdir, name))
except Exception:
pass
continue
out.append((int(prio_s), ts, name))
except (ValueError, IndexError):
continue
except FileNotFoundError:
pass
out.sort()
return out
class _Slot:
"""A held queue slot. Release via .release() or context manager."""
def __init__(self, node, ticket_name, fh, waited):
self.node = node
self.ticket_name = ticket_name
self.fh = fh
self.waited = waited
self._released = False
def release(self):
if self._released:
return
self._released = True
try:
fcntl.flock(self.fh, fcntl.LOCK_UN)
self.fh.close()
except Exception:
pass
# Remove our ticket (best effort — a stale ticket is harmless;
# the next waiter re-reads the directory each poll)
try:
os.unlink(os.path.join(_tickets_dir(self.node), self.ticket_name))
except Exception:
pass
log.debug("cdp_queue: released slot for node=%s (waited %.1fs)",
self.node, self.waited)
def __enter__(self):
return self
def __exit__(self, *exc):
self.release()
def acquire(node, priority=PRIORITY_NORMAL, timeout=ACQUIRE_TIMEOUT):
"""
Block until a CDP slot is free for `node`, then return a _Slot.
Raises QueueTimeout after `timeout` seconds. Raises ValueError for
unknown nodes.
"""
if node not in VALID_NODES:
raise ValueError("unknown node: %r (valid: %s)" % (node, VALID_NODES))
if priority not in (PRIORITY_HIGH, PRIORITY_NORMAL, PRIORITY_LOW):
raise ValueError("invalid priority: %r" % (priority,))
_ensure_dirs(node)
tdir = _tickets_dir(node)
# Our ticket: "<prio>-<timestamp>-<uuid>.ticket", sorted by (prio, ts)
ticket = "%d-%f-%s.ticket" % (priority, time.time(), uuid.uuid4().hex[:8])
open(os.path.join(tdir, ticket), "w").close()
start = time.time()
warned = False
try:
while True:
elapsed = time.time() - start
if elapsed >= timeout:
raise QueueTimeout(
"node=%s: no CDP slot free after %.0fs (priority=%d)" %
(node, timeout, priority))
if elapsed >= WARN_AFTER and not warned:
warned = True
depth = len(_read_tickets(node))
log.warning("cdp_queue: node=%s waiting %.0fs for slot "
"(queue depth %d, priority %d)",
node, elapsed, depth, priority)
tickets = _read_tickets(node)
# Am I among the first MAX_CONCURRENT in line?
# (priority, then arrival time). The first N tickets are all
# eligible to grab slots; they distribute via non-blocking flock.
my_pos = next((i for i, (_, _, name) in enumerate(tickets)
if name == ticket), None)
if my_pos is not None and my_pos < MAX_CONCURRENT:
# Try each slot file non-blocking
for i in range(MAX_CONCURRENT):
fh = open(_slot_path(node, i), "w")
try:
fcntl.flock(fh, fcntl.LOCK_EX | fcntl.LOCK_NB)
except (BlockingIOError, OSError):
fh.close()
continue
# Got it
waited = time.time() - start
if waited > 1.0:
log.debug("cdp_queue: node=%s acquired slot after "
"%.1fs (priority %d)", node, waited, priority)
return _Slot(node, ticket, fh, waited)
time.sleep(POLL_INTERVAL)
except BaseException:
# On timeout or interrupt, remove our ticket so we don't block others
try:
os.unlink(os.path.join(tdir, ticket))
except Exception:
pass
raise
def release(slot):
"""Release a slot returned by acquire()."""
slot.release()
def cdp_slot(node, priority=PRIORITY_NORMAL, timeout=ACQUIRE_TIMEOUT):
"""
Context manager. Usage:
with cdp_slot("opm", priority=PRIORITY_HIGH):
... CDP work ...
"""
return acquire(node, priority=priority, timeout=timeout)
def queue_depth(node):
"""Current number of waiters for a node (for monitoring)."""
if node not in VALID_NODES:
raise ValueError("unknown node: %r" % node)
return len(_read_tickets(node))
-129
View File
@@ -1,129 +0,0 @@
#!/usr/bin/env python3
"""Chat rate metric: messages per minute for main chat and each side chat.
Polls muse-chat-api.py for message counts, calculates delta vs previous poll,
logs rates to a time-series file.
Usage: chat-rate.py --account <name> --cdp-port <port> [--interval 60]
"""
import json, subprocess, sys, time, os, argparse
from datetime import datetime, timezone
API = os.path.expanduser("~/Projects/NetVM/bin/muse-chat-api.py")
STATE_FILE = os.path.expanduser("~/Projects/NetVM/logs/chat-rate-state.json")
LOG_FILE = os.path.expanduser("~/Projects/NetVM/logs/chat-rate.log")
def run_api(account, *args):
"""Run muse-chat-api.py and return stdout."""
cmd = ["python3", API, "--account", account] + list(args)
# Note: cdp-port is baked into the account config, not passed here
result = subprocess.run(cmd, capture_output=True, text=True, timeout=60)
return result.stdout
def count_messages(text):
"""Count messages in API output. Messages are separated by '---'."""
# The API outputs messages separated by ---\n
parts = [p.strip() for p in text.split("---") if p.strip()]
# Filter out non-message lines (headers, etc.)
# Messages typically have substantial content
return len([p for p in parts if len(p) > 10])
def get_side_chats(account):
"""List side chat names."""
out = run_api(account, "sidechat", "list")
# Parse: names and timestamps separated by blank lines
# Format: "Side chats\n\n<name>\n\n<timestamp>\n\n<name>\n\n<timestamp>..."
lines = [l.strip() for l in out.split("\n") if l.strip()]
chats = []
# Skip header "Side chats", then pair up (name, timestamp)
lines = [l for l in lines if l != "Side chats"]
# Lines alternate: name, timestamp, name, timestamp...
for i in range(0, len(lines), 2):
if i < len(lines):
name = lines[i]
# Verify next is a timestamp (ends with m/h/d)
if i + 1 < len(lines) and lines[i+1][-1] in "mhd":
chats.append(name)
return chats
def get_chat_count(account, chat_name=None):
"""Get message count for main or a side chat."""
if chat_name:
# Switch to side chat, get messages, switch back
run_api(account, "sidechat", "use", chat_name)
out = run_api(account, "messages")
run_api(account, "sidechat", "main") # switch back
else:
out = run_api(account, "messages")
return count_messages(out)
def main():
p = argparse.ArgumentParser()
p.add_argument("--account", required=True)
p.add_argument("--interval", type=int, default=60, help="poll interval seconds")
p.add_argument("--once", action="store_true", help="single poll, no loop")
args = p.parse_args()
# Load previous state
prev = {}
if os.path.exists(STATE_FILE):
with open(STATE_FILE) as f:
prev = json.load(f)
def poll():
now = datetime.now(timezone.utc).isoformat()
counts = {}
# Main chat
try:
counts["main"] = get_chat_count(args.account)
except Exception as e:
print(f"main: error {e}", file=sys.stderr)
# Side chats
try:
sc_list = get_side_chats(args.account)
print(f"DEBUG: found {len(sc_list)} side chats", file=sys.stderr)
for sc in sc_list:
try:
counts[f"side:{sc}"] = get_chat_count(args.account, sc)
except Exception as e:
print(f"side:{sc}: error {e}", file=sys.stderr)
except Exception as e:
print(f"sidechat list: error {e}", file=sys.stderr)
# Calculate rates
results = []
for chat, count in counts.items():
rate = 0.0
if chat in prev:
prev_count, prev_time = prev[chat]
dt = (datetime.fromisoformat(now) - datetime.fromisoformat(prev_time)).total_seconds() / 60.0
if dt > 0:
rate = (count - prev_count) / dt
results.append((now, chat, count, round(rate, 2)))
prev[chat] = (count, now)
# Log
os.makedirs(os.path.dirname(LOG_FILE), exist_ok=True)
with open(LOG_FILE, "a") as f:
for ts, chat, count, rate in results:
f.write(f"{ts} {chat} count={count} rate={rate}/min\n")
# Save state
with open(STATE_FILE, "w") as f:
json.dump(prev, f)
# Print
for ts, chat, count, rate in results:
print(f"{chat}: {count} msgs, {rate}/min")
if args.once:
poll()
else:
while True:
poll()
time.sleep(args.interval)
if __name__ == "__main__":
main()
-281
View File
@@ -1,281 +0,0 @@
#!/usr/bin/env python3
"""chat-state-check.py — one-shot CDP chat-state probe for a single node.
Runs INSIDE the node's netns (CDP listens on 127.0.0.1 there).
Usage: chat-state-check.py <cdp_port> <mode> [name]
modes: main | list | sidechat <name>
Prints one JSON object to stdout.
Honest limits (see CHATSTATE_SPEC.md): only the DOM-visible message window
is captured (React virtualization); author attribution is best-effort and
may be null; checked_at is the report time, not per-message times.
"""
import base64
import json
import sys
import time
import urllib.request
import websocket
MSG_N = 10
MSG_WIDTH = 500
def connect(port):
with urllib.request.urlopen(
"http://127.0.0.1:%s/json/list" % port, timeout=5) as r:
targets = json.load(r)
pages = [t for t in targets if t.get("type") == "page"]
if not pages:
return None
return websocket.create_connection(
pages[0]["webSocketDebuggerUrl"], timeout=15)
def ev1(ws, expr, await_p=False, reads=30):
"""Runtime.evaluate that skips CDP event chatter while awaiting ours."""
ws.send(json.dumps({
"id": 1, "method": "Runtime.evaluate",
"params": {"expression": expr, "returnByValue": True,
"awaitPromise": await_p}}))
for _ in range(reads):
resp = json.loads(ws.recv())
if resp.get("id") != 1:
continue
if "error" in resp:
raise RuntimeError("CDP evaluate failed: %s" % resp["error"])
res = resp.get("result", {}).get("result", {})
return res.get("value")
raise RuntimeError("CDP evaluate: no response")
def read_current(ws):
title = ev1(ws, "document.title") or ""
url = ev1(ws, "window.location.href") or ""
msgs = ev1(ws, """(() => {
const ps = [...document.querySelectorAll('p')].slice(-%d)
.map(p => (p.innerText||'').slice(0,%d)).filter(t => t.trim());
return ps;
})()""" % (MSG_N, MSG_WIDTH)) or []
return {"title": title, "url": url,
"messages": [{"text": t} for t in msgs]}
def _wait_for(ws, expr, timeout_s=12):
"""Poll a JS truthiness expression until true or timeout."""
deadline = time.time() + timeout_s
while time.time() < deadline:
try:
if ev1(ws, expr):
return True
except Exception: # noqa: BLE001
pass
time.sleep(1)
return False
def ensure_sidebar_open(ws):
"""The toggle closes an open sidebar — only click when closed.
Detected structurally via the 'New side chat' button (text matching
'Side chats' is unreliable: main-chat messages can contain the phrase).
Waits for React to render the panel after opening.
"""
is_open = ev1(ws, """(() => {
return !!document.querySelector('button[aria-label="New side chat"]');
})()""")
if not is_open:
ev1(ws, """(() => {
const btn = [...document.querySelectorAll('button')].find(b =>
(b.textContent||'').includes('Open chat and side chats'));
if (btn) btn.click();
})()""")
_wait_for(ws, """(() => {
return !!document.querySelector(
'button[aria-label="New side chat"]');
})()""")
# The row list renders a beat after the panel; wait for it.
_wait_for(ws, """(() => {
return document.querySelectorAll(
'button[aria-label="More thread actions"]').length > 0;
})()""", timeout_s=10)
_SIDEBAR_JS = """(() => {
const anchor = document.querySelector('button[aria-label="New side chat"]');
if (!anchor) return 'NOANCHOR';
let sec = anchor.parentElement;
for (let i = 0; i < 8 && sec; i++) {
if (sec.querySelectorAll(
'button[aria-label="More thread actions"]').length) break;
sec = sec.parentElement;
}
if (!sec) return 'NOSEC';
return JSON.stringify({ok: true});
})()"""
def _sidebar_section(ws):
"""Return True when the side-chat list section is addressable."""
raw = ev1(ws, _SIDEBAR_JS)
try:
return json.loads(raw or "").get("ok", False)
except (ValueError, TypeError):
return False
_LIST_JS = """(() => {
const T = (el) => (el.innerText || '').trim();
const TS =
/^(\\d+[smhd]|just now|Yesterday|Today|[A-Z][a-z]{2} \\d{1,2}.*)$/;
const Y = (el) => el.getBoundingClientRect().top;
const h3s = [...document.querySelectorAll('h3')];
const head = h3s.find(h => h.textContent.trim() === 'Side chats');
if (!head) return '[]';
// Scroll the list to top so the section's rows sit above the next
// (sticky) header; without this, rows below the fold are missed.
let sc = head.parentElement;
for (let i = 0; i < 8 && sc; i++) {
if (sc.scrollHeight > sc.clientHeight + 10) break;
sc = sc.parentElement;
}
if (sc) sc.scrollTop = 0;
const y0 = Y(head);
const isStopText = (t) => t === 'Unread updates' || t === 'Chats' ||
t === 'Unread chats';
let stopY = null;
for (const h of h3s) {
const y = Y(h);
if (y > y0 + 5 && (stopY === null || y < stopY)) stopY = y;
}
// Non-H3 section headers ("Unread updates", ...) also bound the list.
for (const el of document.querySelectorAll('*')) {
if (el.children.length !== 0) continue;
if (!isStopText((el.textContent || '').trim())) continue;
const y = Y(el);
if (y > y0 + 5 && (stopY === null || y < stopY)) stopY = y;
}
const names = [];
for (const el of document.querySelectorAll('div')) {
const cn = (el.className || '').toString();
if (cn.indexOf('nav-row') === -1) continue;
const y = Y(el);
if (y <= y0 + 5) continue;
if (stopY !== null && y >= stopY - 5) continue;
const m = T(el).match(/^([^\\n]+)\\n\\n(.+)$/);
if (m && TS.test(m[2].trim()) && names.indexOf(m[1].trim()) === -1) {
names.push(m[1].trim());
if (names.length >= 8) break;
}
}
return JSON.stringify(names);
})()"""
def _list_once(ws):
raw = ev1(ws, _LIST_JS)
try:
return json.loads(raw or "[]")
except (ValueError, TypeError):
return []
def list_sidechats(ws):
ensure_sidebar_open(ws)
if not _sidebar_section(ws):
return []
# React renders rows progressively after navigation; poll until the
# list stabilizes instead of trusting the first paint.
prev = None
for _ in range(4):
names = _list_once(ws)
if names and names == prev:
return names
prev = names
time.sleep(2)
return prev or []
def open_sidechat(ws, name):
"""Open the side chat by clicking its row div inside the sidebar section.
Clicking the row (not a text search over the whole document) avoids
hitting message text that happens to contain the chat name.
"""
ensure_sidebar_open(ws)
b64 = base64.b64encode(name.encode()).decode()
return ev1(ws, """(async () => {
const nm = atob('%s');
const T = (el) => (el.innerText || '').trim();
const target = [...document.querySelectorAll('div')].find(el => {
const cn = (el.className || '').toString();
if (cn.indexOf('nav-row') === -1) return false;
const m = T(el).match(/^([^\\n]+)\\n\\n(.+)$/);
return m && m[1].trim() === nm;
});
if (!target) return 'NOTFOUND';
target.click();
await new Promise(r => setTimeout(r, 4000));
const href = window.location.href;
if (href.indexOf('/thread/') === -1) return 'NOTFOUND';
return href;
})()""" % b64, await_p=True)
def navigate_main(ws):
ws.send(json.dumps({"id": 2, "method": "Page.navigate",
"params": {"url": "https://muse.ai/"}}))
for _ in range(20):
resp = json.loads(ws.recv())
if resp.get("id") == 2:
break
time.sleep(6)
def main():
if len(sys.argv) < 3:
print(json.dumps({"ok": False, "error": "usage"}))
sys.exit(1)
port, mode = sys.argv[1], sys.argv[2]
name = sys.argv[3] if len(sys.argv) > 3 else ""
try:
ws = connect(port)
except Exception as e: # noqa: BLE001
print(json.dumps({"ok": False, "error": "cdp_connect_failed",
"detail": str(e)[:120]}))
return
if ws is None:
print(json.dumps({"ok": False, "error": "no_page"}))
return
try:
if mode == "list":
print(json.dumps({"ok": True,
"sidechats": list_sidechats(ws)}))
elif mode == "sidechat":
href = open_sidechat(ws, name)
if href == "NOTFOUND":
print(json.dumps({"ok": False, "error": "sidechat_not_found",
"name": name}))
else:
cur = read_current(ws)
cur.update({"ok": True, "context": "sidechat", "name": name})
print(json.dumps(cur))
else:
navigate_main(ws)
cur = read_current(ws)
cur.update({"ok": True, "context": "main", "name": "Main chat"})
print(json.dumps(cur))
except Exception as e: # noqa: BLE001
print(json.dumps({"ok": False, "error": "probe_failed",
"detail": str(e)[:200]}))
finally:
try:
ws.close()
except Exception: # noqa: BLE001
pass
main()
-129
View File
@@ -1,129 +0,0 @@
#!/usr/bin/env python3
"""chat-state-report.py — per-node chat-state tap for box.
Probes each NetVM node's browser via CDP (inside its netns): main chat +
side chats, last messages. Signs and POSTs to the board chat-state ingest,
mirroring accounts-health.sh: payload is
<machine>\\n<ts>\\n<facts-json>, namespace "health".
Usage: chat-state-report.py [--no-post]
Cron (on bl, every 5 min):
*/5 * * * * ~/Projects/NetVM/bin/chat-state-report.py >/dev/null 2>&1
Health key setup: same key as accounts-health (namespace "health",
registered in /srv/board/health_signers for machine bl).
"""
import importlib.util
import json
import os
import subprocess
import sys
import tempfile
import time
NETVM_DIR = os.environ.get("NETVM_DIR",
os.path.expanduser("~/Projects/NetVM"))
CHECK = os.path.join(NETVM_DIR, "bin", "chat-state-check.py")
MACHINE = os.environ.get("MUSE_MACHINE", "bl")
KEY = os.environ.get("CHATSTATE_KEY",
os.path.expanduser("~/.ssh/muse-health"))
ENDPOINT = os.environ.get(
"CHATSTATE_ENDPOINT",
"https://board.muse-dev.online/api/box/chat-state/report")
MAX_SIDECHATS = 5
FACTS_CAP = 64 * 1024
POST = "--no-post" not in sys.argv
def load_nodes():
spec = importlib.util.spec_from_file_location(
"netvm_registry",
os.path.join(NETVM_DIR, "bin", "netvm-registry.py"))
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return mod.load() # node -> {"cdp_port": ...}
def run_check(node, port, *args):
cmd = ["sudo", "-n", "ip", "netns", "exec", "warp-%s" % node,
"python3", CHECK, str(port)] + list(args)
try:
r = subprocess.run(cmd, capture_output=True, text=True, timeout=150)
line = r.stdout.strip().splitlines()[-1]
return json.loads(line)
except Exception as e: # noqa: BLE001
return {"ok": False, "error": "check_failed",
"detail": str(e)[:120]}
def main():
nodes = load_nodes()
chats = []
now = int(time.time())
for node, rec in sorted(nodes.items()):
port = rec.get("cdp_port")
if not port:
continue
base = {"node": node, "checked_at": now}
m = run_check(node, port, "main")
entry = dict(base)
if m.get("ok"):
entry.update(m)
else:
entry.update({"context": "main", "error": m.get("error"),
"detail": m.get("detail", "")[:120]})
chats.append(entry)
continue
chats.append(entry)
lst = run_check(node, port, "list")
for name in (lst.get("sidechats") or [])[:MAX_SIDECHATS]:
s = run_check(node, port, "sidechat", name)
sentry = dict(base)
if s.get("ok"):
sentry.update(s)
else:
sentry.update({"context": "sidechat", "name": name,
"error": s.get("error"),
"detail": s.get("detail", "")[:120]})
chats.append(sentry)
facts = {"chats": chats, "checked_at": now}
facts_json = json.dumps(facts, separators=(",", ":"))
if len(facts_json) > FACTS_CAP:
# Shed message bodies, keep metadata.
for c in chats:
c.pop("messages", None)
facts["truncated"] = True
facts_json = json.dumps(facts, separators=(",", ":"))
if not POST:
print(json.dumps(json.loads(facts_json), indent=2))
return
if not os.path.exists(KEY):
print("chat-state-report: %s missing — printing JSON, not posting"
% KEY, file=sys.stderr)
print(facts_json)
return
tmp = tempfile.mkdtemp()
try:
payload = os.path.join(tmp, "payload")
with open(payload, "w") as f:
f.write("%s\n%d\n%s" % (MACHINE, now, facts_json))
# Fresh temp dir: no stale .sig to worry about (cf. AGENTS.md).
subprocess.run(["ssh-keygen", "-Y", "sign", "-f", KEY,
"-n", "health", payload],
check=True, capture_output=True)
with open(payload + ".sig") as f:
sig = f.read()
body = {"machine": MACHINE, "ts": now, "facts_json": facts_json,
"signature": sig}
r = subprocess.run(
["curl", "-s", "-X", "POST", ENDPOINT,
"-H", "Content-Type: application/json",
"--data", json.dumps(body)],
capture_output=True, text=True, timeout=90)
print(r.stdout[:300])
finally:
subprocess.run(["rm", "-rf", tmp])
main()
-99
View File
@@ -1,99 +0,0 @@
#!/bin/bash
# chrome-error-scan.sh - scan per-profile chrome logs for concerning patterns.
# Self-contained: scans, compares against watermark, reports only NEW matches.
#
# Usage: chrome-error-scan.sh [--json] [--no-advance]
# Default output: "profile:new_count" lines for profiles with new matches,
# or "OK: no new errors" if clean.
# --json: output JSON {"profile": {"total": N, "new": M}, ...}
# --no-advance: report against the watermark WITHOUT advancing it.
# Peek-only read for high-frequency pollers (e.g. the web
# surface via `box-ctl chrome-errors --no-advance`). Runs
# without the flag keep the classic advance-on-read semantics.
#
# Watermark: /home/super/Projects/NetVM/chrome-error-watermark.json
# Patterns: FATAL, crash, segfault, out of memory (case-insensitive)
LOGDIR="/home/super/Projects/NetVM"
WATERMARK="$LOGDIR/chrome-error-watermark.json"
AS_JSON=0
NO_ADVANCE=0
for arg in "$@"; do
case "$arg" in
--json) AS_JSON=1 ;;
--no-advance) NO_ADVANCE=1 ;;
*) echo "chrome-error-scan.sh: unknown argument: $arg" >&2; exit 2 ;;
esac
done
# Gather current counts per profile (grep -c prints 0 with exit 1 on no match;
# the || true masks the exit code while preserving the "0" on stdout)
get_count() {
local f="$LOGDIR/chromebox-$1.log"
if [ -f "$f" ]; then
grep -ciE "FATAL|crash|segfault|out of memory" "$f" 2>/dev/null || true
else
echo 0
fi
}
MUSE_C=$(get_count muse)
PIP_C=$(get_count pip)
N646_C=$(get_count 646)
OPM_C=$(get_count opm)
python3 - "$WATERMARK" "$MUSE_C" "$PIP_C" "$N646_C" "$OPM_C" "$AS_JSON" "$NO_ADVANCE" <<'PYEOF'
import json, sys, os
watermark_path = sys.argv[1]
current = {
"muse": int(sys.argv[2]),
"pip": int(sys.argv[3]),
"646": int(sys.argv[4]),
"opm": int(sys.argv[5]),
}
as_json = len(sys.argv) > 6 and sys.argv[6] == "1"
no_advance = len(sys.argv) > 7 and sys.argv[7] == "1"
# Load watermark (tolerate missing/corrupt file -> treat as all-zero)
watermark = {}
if os.path.exists(watermark_path):
try:
with open(watermark_path) as f:
data = json.load(f)
watermark = data.get("counts", data) if isinstance(data, dict) else {}
except (ValueError, IOError, OSError):
watermark = {}
new_counts = {}
for prof in ("muse", "pip", "646", "opm"):
old = watermark.get(prof, 0)
try:
old = int(old)
except (TypeError, ValueError):
old = 0
delta = current[prof] - old
new_counts[prof] = max(delta, 0)
if as_json:
out = {p: {"total": current[p], "new": new_counts[p]} for p in current}
print(json.dumps(out))
else:
any_new = False
for prof in ("muse", "pip", "646", "opm"):
if new_counts[prof] > 0:
print("%s:%d" % (prof, new_counts[prof]))
any_new = True
if not any_new:
print("OK: no new errors")
# Update watermark atomically (skipped in --no-advance peek mode: the
# caller gets a read-only view and the CLI's advance-on-read semantics are
# left untouched).
if not no_advance:
tmp = watermark_path + ".tmp"
with open(tmp, "w") as f:
json.dump({"counts": current}, f, indent=2)
os.replace(tmp, watermark_path)
PYEOF
-364
View File
@@ -1,364 +0,0 @@
#!/usr/bin/env python3
"""
chromebox-gateway.py — HTTPS fallback for chromebox/DM control when SSH is down.
The chromeboxes (headless Chromium per NetVM node) are normally driven over
SSH via muse-chat-api.py / dm.py, which speak CDP through the per-node
netns relay. When SSH to bl breaks, this gateway provides a constrained
HTTPS path to the same high-level operations.
CRITICAL: this does NOT expose raw CDP. Raw CDP (Runtime.evaluate,
Page.navigate, etc.) is arbitrary code execution inside the browser with
the agent's live session. This gateway exposes ONLY the allowlisted
high-level operations below, each mapped to an existing audited script.
Usage:
python3 chromebox-gateway.py --port 8444
Auth: Bearer token, reusing the shared exec per-agent token files
(~/.exec-tokens/<agent>). The master token (~/.exec-server-token) also
works. Identity = token filename; tokens are never logged.
Endpoints:
GET /health {"status":"ok"} — no auth (load-balancer friendly)
POST /api/v1/op {"op": "<name>", "params": {...}} — bearer auth
Audit: every call appended to ~/.chromebox-gateway-audit.jsonl (0600).
"""
import argparse
import hmac
import json
import os
import re
import ssl
import subprocess
import sys
import time
from http.server import HTTPServer, BaseHTTPRequestHandler
from urllib.parse import urlparse
# ---------------------------------------------------------------- config
BIN_DIR = os.path.expanduser("~/Projects/NetVM/bin")
TOKEN_FILE = "/home/super/.exec-server-token" # master token (shared exec token file)
TOKEN_DIR = "/home/super/.exec-tokens" # per-agent tokens (shared exec token files)
AUDIT_FILE = "/home/super/.chromebox-gateway-audit.jsonl"
CERT_FILE = "/home/super/.chromebox-gateway-cert.pem"
KEY_FILE = "/home/super/.chromebox-gateway-key.pem"
NODES = ("muse", "pip", "646", "opm")
MAX_MSG = 2000 # gateway-level cap; downstream scripts enforce their own
BACKEND_TIMEOUT = 120 # seconds per backend call
# Rate limit: token bucket per identity — 20 req/min sustained, burst 5.
RATE_PER_SEC = 20.0 / 60.0
RATE_BURST = 5
# ---------------------------------------------------------------- ops
# Each op maps to an argv builder for an existing script. No shell=True,
# ever. Params are validated before building argv.
def _node(params):
node = params.get("account") or params.get("agent")
if node not in NODES:
raise ValueError("account/agent must be one of %s" % (",".join(NODES)))
return node
def _msg(params):
m = params.get("message", "")
if not isinstance(m, str) or not m.strip():
raise ValueError("message must be a non-empty string")
if len(m) > MAX_MSG:
raise ValueError("message exceeds %d chars" % MAX_MSG)
return m
def _target(params):
t = params.get("target", "main")
if not isinstance(t, str) or not t or len(t) > 128:
raise ValueError("target must be a short string")
if not re.fullmatch(r"[A-Za-z0-9_./:-]+", t):
raise ValueError("target has invalid characters")
if ".." in t:
raise ValueError("target must not contain '..'")
return t
def _n(params, default=5, cap=50):
n = params.get("n", default)
try:
n = int(n)
except (TypeError, ValueError):
raise ValueError("n must be an integer")
if not 1 <= n <= cap:
raise ValueError("n must be 1..%d" % cap)
return n
def _tags(params):
tags = params.get("tags", [])
if not isinstance(tags, list):
raise ValueError("tags must be a list")
out = []
for t in tags:
if not isinstance(t, str) or len(t) > 128:
raise ValueError("bad tag")
if not re.fullmatch(r"[A-Za-z0-9_:=\-./]+", t):
raise ValueError("tag has invalid characters: %r" % t[:40])
out.append(t)
if len(out) > 10:
raise ValueError("too many tags (max 10)")
return out
CHAT = os.path.join(BIN_DIR, "muse-chat-api.py")
DM = os.path.join(BIN_DIR, "dm.py")
def op_chat_send(p):
return [CHAT, "--account", _node(p), "send", _msg(p)]
def op_chat_messages(p):
return [CHAT, "--account", _node(p), "messages", "--n", str(_n(p))]
def op_chat_sidechats(p):
return [CHAT, "--account", _node(p), "sidechat", "list"]
def op_chat_sidechat_create(p):
argv = [CHAT, "--account", _node(p), "sidechat", "create"]
name = p.get("name")
if name:
if not isinstance(name, str) or len(name) > 80 or not re.fullmatch(r"[A-Za-z0-9 _-]+", name):
raise ValueError("bad sidechat name")
argv += ["--name", name]
return argv
def op_chat_approvals(p):
return [CHAT, "--account", _node(p), "approvals"]
def op_chat_url(p):
return [CHAT, "--account", _node(p), "url"]
def op_dm_send(p):
node = _node(p)
to = p.get("to", node)
if to not in NODES:
raise ValueError("to must be one of %s" % (",".join(NODES)))
argv = [DM, "send", "--agent", node, "--to", to,
"--target", _target(p), _msg(p)]
for t in _tags(p):
argv += ["--tag", t]
return argv
def op_dm_read(p):
return [DM, "read", "--agent", _node(p),
"--target", _target(p), "--n", str(_n(p, default=5, cap=20))]
# The allowlist. Adding an op here is a security decision — review accordingly.
# Deliberately absent: wait (long-poll), upload (file ingress),
# sidechat use (raw navigation), anything raw-CDP.
ALLOWLIST = {
"chat.send": ("write", op_chat_send),
"chat.messages": ("read", op_chat_messages),
"chat.sidechats": ("read", op_chat_sidechats),
"chat.sidechat_create": ("write", op_chat_sidechat_create),
"chat.approvals": ("read", op_chat_approvals),
"chat.url": ("read", op_chat_url),
"dm.send": ("write", op_dm_send),
"dm.read": ("read", op_dm_read),
}
# ---------------------------------------------------------------- auth
def _read_token_file(path):
try:
with open(path) as f:
return f.read().strip()
except OSError:
return ""
def check_token(token):
"""Return identity label or None. Tokens never leave this function."""
if not token:
return None
if hmac.compare_digest(token, _read_token_file(TOKEN_FILE)):
return "master"
try:
names = os.listdir(TOKEN_DIR)
except OSError:
return None
for name in names:
if not re.fullmatch(r"[A-Za-z0-9_-]+", name):
continue
t = _read_token_file(os.path.join(TOKEN_DIR, name))
if t and hmac.compare_digest(token, t):
return name
return None
# ---------------------------------------------------------------- rate limit
_buckets = {} # identity -> [tokens, last_ts]
def rate_ok(identity):
now = time.monotonic()
tokens, last = _buckets.get(identity, (RATE_BURST, now))
tokens = min(RATE_BURST, tokens + (now - last) * RATE_PER_SEC)
if tokens < 1.0:
_buckets[identity] = (tokens, now)
return False
_buckets[identity] = (tokens - 1.0, now)
return True
# ---------------------------------------------------------------- audit
def audit(entry):
entry = dict(entry)
entry["ts"] = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
# Defense in depth: never let a full message body reach the audit file,
# even if a caller forgets to summarize first.
params = entry.get("params")
if isinstance(params, dict) and "message" in params:
params = dict(params)
params["message_len"] = len(params.pop("message"))
entry["params"] = params
try:
fd = os.open(AUDIT_FILE, os.O_WRONLY | os.O_CREAT | os.O_APPEND, 0o600)
with os.fdopen(fd, "a") as f:
f.write(json.dumps(entry) + "\n")
except OSError as e:
print("audit write failed: %s" % e, file=sys.stderr)
# ---------------------------------------------------------------- handler
class Handler(BaseHTTPRequestHandler):
server_version = "chromebox-gateway/1.0"
def log_message(self, fmt, *args): # quiet; audit log is the record
pass
def _json(self, code, obj):
body = json.dumps(obj).encode()
self.send_response(code)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(body)))
self.end_headers()
self.wfile.write(body)
def do_GET(self):
if urlparse(self.path).path == "/health":
self._json(200, {"status": "ok", "ops": sorted(ALLOWLIST)})
return
self._json(404, {"error": "not_found"})
def do_POST(self):
if urlparse(self.path).path != "/api/v1/op":
self._json(404, {"error": "not_found"})
return
auth = self.headers.get("Authorization", "")
token = auth[7:] if auth.startswith("Bearer ") else ""
identity = check_token(token)
if not identity:
self._json(401, {"error": "unauthorized"})
return
if not rate_ok(identity):
audit({"identity": identity, "op": None, "result": "rate_limited"})
self._json(429, {"error": "rate_limited"})
return
try:
length = int(self.headers.get("Content-Length", 0))
except ValueError:
length = 0
if length > 65536:
self._json(413, {"error": "body_too_large"})
return
try:
req = json.loads(self.rfile.read(length) or b"{}")
except (ValueError, OSError):
self._json(400, {"error": "bad_json"})
return
op = req.get("op")
params = req.get("params") or {}
if not isinstance(params, dict):
self._json(400, {"error": "params_must_be_object"})
return
entry = ALLOWLIST.get(op)
if not entry:
audit({"identity": identity, "op": op, "result": "unknown_op"})
self._json(400, {"error": "unknown_op", "allowed": sorted(ALLOWLIST)})
return
cls, builder = entry
t0 = time.monotonic()
try:
argv = builder(params)
except ValueError as e:
audit({"identity": identity, "op": op, "class": cls, "result": "bad_params",
"detail": str(e)[:120]})
self._json(400, {"error": "bad_params", "detail": str(e)})
return
# Redacted summary for the audit log — never the full message.
summary = {k: (v[:80] + "…" if isinstance(v, str) and len(v) > 80 else v)
for k, v in params.items() if k != "message"}
if "message" in params:
summary["message_len"] = len(params["message"])
try:
proc = subprocess.run(argv, capture_output=True, text=True,
timeout=BACKEND_TIMEOUT)
ok = proc.returncode == 0
result = "ok" if ok else "backend_error"
self._json(200 if ok else 502, {
"ok": ok,
"op": op,
"returncode": proc.returncode,
"stdout": proc.stdout[-8000:],
"stderr": proc.stderr[-2000:],
})
except subprocess.TimeoutExpired:
result = "timeout"
self._json(504, {"ok": False, "op": op, "error": "backend_timeout"})
except OSError as e:
result = "exec_failed"
self._json(500, {"ok": False, "op": op, "error": "exec_failed"})
finally:
audit({"identity": identity, "op": op, "class": cls,
"node": params.get("account") or params.get("agent"),
"params": summary, "result": result,
"latency_ms": int((time.monotonic() - t0) * 1000)})
# ---------------------------------------------------------------- main
def ensure_cert():
if os.path.exists(CERT_FILE) and os.path.exists(KEY_FILE):
return
print("generating self-signed cert...", file=sys.stderr)
subprocess.run([
"openssl", "req", "-x509", "-newkey", "rsa:2048",
"-keyout", KEY_FILE, "-out", CERT_FILE,
"-days", "825", "-nodes", "-subj", "/CN=chromebox-gateway",
], check=True)
os.chmod(KEY_FILE, 0o600)
os.chmod(CERT_FILE, 0o600)
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--port", type=int, default=8444)
ap.add_argument("--bind", default="100.123.153.75",
help="tailnet IP; use 127.0.0.1 for local-only")
args = ap.parse_args()
ensure_cert()
context = ssl.SSLContext(ssl.PROTOCOL_TLS_SERVER)
context.load_cert_chain(CERT_FILE, KEY_FILE)
srv = HTTPServer((args.bind, args.port), Handler)
srv.socket = context.wrap_socket(srv.socket, server_side=True)
print("chromebox-gateway listening on https://%s:%d" % (args.bind, args.port),
file=sys.stderr)
print("ops: %s" % ", ".join(sorted(ALLOWLIST)), file=sys.stderr)
try:
srv.serve_forever()
except KeyboardInterrupt:
pass
if __name__ == "__main__":
main()
-225
View File
@@ -1,225 +0,0 @@
#!/usr/bin/env bash
# chromebox-watchdog.sh [profile] — keep a chrome-box profile alive and healthy.
# Checks: 1) chromium process for the profile is running,
# 2) CDP responds and the chat page is present.
# If unhealthy: kill any stale chrome for the profile and relaunch headless
# via netvm-chrome.sh (profile dir persists session/cookies — the process is
# disposable, the state is not). Mirrors operator-646's container
# recover-after-rebuild.sh philosophy.
# Runs every 2 min via systemd timer chromebox-watchdog-<profile>.timer.
set -euo pipefail
export XDG_RUNTIME_DIR="${XDG_RUNTIME_DIR:-/run/user/$(id -u)}"
export DBUS_SESSION_BUS_ADDRESS="${DBUS_SESSION_BUS_ADDRESS:-unix:path=${XDG_RUNTIME_DIR}/bus}"
# Prevent overlapping runs: the timer fires every 2 min but a relaunch
# (kill + sleep 25 + chrome startup + page load) can exceed that, and two
# concurrent runs kill each others chrome (observed 2026-10-03: pip flapped
# with simultaneous "relaunch OK" and "relaunch FAILED").
LOCK="/tmp/chromebox-watchdog-${1:-pip}.lock"
# Tests source this file with CHROMEBOX_WATCHDOG_LIB_ONLY=1: they resolve
# ports and call helpers without running checks, so no lock is needed.
if [ "${CHROMEBOX_WATCHDOG_LIB_ONLY:-}" != "1" ]; then
exec 9>"$LOCK"
if ! flock -n 9; then
echo "[$(date -u +%FT%TZ)] [$1] another watchdog run in progress, skipping" >&2
exit 0
fi
fi
PROFILE="${1:-pip}"
NETVM_BIN="/home/super/Projects/NetVM/bin"
LOG="/home/super/Projects/NetVM/chromebox-watchdog.log"
CHROME_LOG="/home/super/Projects/NetVM/chromebox-${PROFILE}.log"
# Rotate a log file past 10MB (keep one generation)
rotate_log() {
local f="$1"
[ -f "$f" ] || return 0
local sz
sz=$(stat -c%s "$f" 2>/dev/null || echo 0)
if [ "$sz" -gt 10485760 ]; then
mv -f "$f" "$f.1"
echo "[$(date -u +%FT%TZ)] [$PROFILE] log rotated" > "$f"
fi
}
rotate_log "$LOG"
# CHROME_LOG rotation happens after PROFILE is set (see below)
# Ports come from the fleet registry, not a hardcoded list: every active
# node (def/dev included) gets supervision automatically. The old 4-profile
# case left dev/def unsupervised — a dead Warp tunnel paged forever with
# no auto-recovery (2026-10-06 dev outage).
CDP_PORT="$("$NETVM_BIN/netvm-registry.py" "$PROFILE" 2>/dev/null)" || {
echo "unknown profile: $PROFILE" >&2
exit 1
}
[ -n "$CDP_PORT" ] || { echo "unknown profile: $PROFILE" >&2; exit 1; }
rotate_log "$CHROME_LOG"
log() { echo "[$(date -u +%FT%TZ)] [$PROFILE] $*" | tee -a "$LOG"; }
cdp_list() {
"$NETVM_BIN/netvm-exec.sh" "$PROFILE" -- curl -s -m 8 "http://127.0.0.1:$CDP_PORT/json/list" 2>/dev/null
}
# HEALTH_FAIL_REASON is set by healthy() on failure: which stage broke.
HEALTH_FAIL_REASON=""
healthy() {
HEALTH_FAIL_REASON=""
pgrep -f "chromium.*profiles/${PROFILE}/" >/dev/null 2>&1 \
|| { HEALTH_FAIL_REASON="no chromium process for profile"; return 1; }
local list
list="$(cdp_list)" \
|| { HEALTH_FAIL_REASON="CDP unreachable on :$CDP_PORT"; return 1; }
[[ "$list" == *'"type": "page"'* ]] \
|| { HEALTH_FAIL_REASON="CDP up but no page target in list"; return 1; }
[[ "$list" == *'"url": "https://muse.ai'* ]] \
|| { HEALTH_FAIL_REASON="CDP up but not on muse.ai"; return 1; }
warp_egress_healthy || return 1
return 0
}
# Stage 5: Warp egress health (2026-10-04). A partitioned browser (tunnel down,
# CDP green) passes stages 1-4 while being unable to reach muse.ai. Check the
# WireGuard handshake age and probe egress from inside the node's netns.
# Sets HEALTH_FAIL_REASON with a distinct "warp egress down" prefix so
# root-cause analysis can distinguish partitions from Chromium crashes.
warp_egress_healthy() {
local wg_out hs_line age_s code
wg_out="$(sudo -n ip netns exec "warp-${PROFILE}" wg show 2>/dev/null)" \
|| { HEALTH_FAIL_REASON="warp egress down (wg show failed in warp-${PROFILE})"; return 1; }
hs_line="$(printf '%s\n' "$wg_out" | grep -i "latest handshake" | head -1)"
[ -n "$hs_line" ] \
|| { HEALTH_FAIL_REASON="warp egress down (no WireGuard handshake in warp-${PROFILE})"; return 1; }
# "latest handshake: 1 minute, 41 seconds ago" -> total seconds
age_s="$(printf '%s\n' "$hs_line" | python3 -c '
import sys, re
s = sys.stdin.read()
m = re.search(r"(\d+)\s*hour", s); h = int(m.group(1)) if m else 0
m = re.search(r"(\d+)\s*minute", s); mi = int(m.group(1)) if m else 0
m = re.search(r"(\d+)\s*second", s); se = int(m.group(1)) if m else 0
print(h*3600 + mi*60 + se)
' 2>/dev/null)"
{ [ -n "$age_s" ] && [ "$age_s" -ge 0 ]; } 2>/dev/null \
|| { HEALTH_FAIL_REASON="warp egress down (unparseable handshake: $hs_line)"; return 1; }
[ "$age_s" -le 180 ] \
|| { HEALTH_FAIL_REASON="warp egress down (handshake ${age_s}s old in warp-${PROFILE})"; return 1; }
code="$(sudo -n ip netns exec "warp-${PROFILE}" curl -m 5 -s -o /dev/null -w "%{http_code}" "https://muse.ai/" 2>/dev/null)" \
|| { HEALTH_FAIL_REASON="warp egress down (egress probe curl failed in warp-${PROFILE})"; return 1; }
# Any 2xx/3xx means we reached muse.ai infra ("/" 307-redirects to the
# auth flow). The probe tests egress connectivity, not page content.
case "$code" in
2*|3*) return 0 ;;
*) HEALTH_FAIL_REASON="warp egress down (egress probe HTTP $code in warp-${PROFILE})"; return 1 ;;
esac
}
# Allow sourcing for tests without running checks.
if [ "${CHROMEBOX_WATCHDOG_LIB_ONLY:-}" = "1" ]; then
return 0 2>/dev/null || exit 0
fi
if healthy; then
# Auto-reconcile idle workers for healthy profiles
# DISABLED 2026-10-06 by operator-646: kpi auto-spawn ignores job schedule fields;
# find_pending_work_for_node returns the alphabetically-first definition every tick,
# re-spawning and re-noticing every ~2min (pip b01 BOX-AUTO-WORKER loop, 29+ copies).
# Watchdog health path untouched. Re-enable once the spawner is schedule-aware.
# python3 "$NETVM_BIN/super-cli.py" kpi auto-spawn --node "$PROFILE" >>"$LOG" 2>&1 || true
exit 0
fi
# Relaunch-loop guard (2026-10-04): if the main browser for this profile
# launched <2 min ago it's probably still starting up (CDP not yet bound).
# Relaunching now would kill a healthy-but-slow cold start via netvm-chrome.sh
# and reset the startup clock every cycle (observed: opm piled up 5 chromiums
# because the guard only protected the kill step, not the relaunch). Skip the
# entire cycle instead.
# NOTE: match only the main browser process (--remote-debugging-port present,
# no --type= flag). Renderer/gpu children (--type=renderer etc.) start later
# than the main process and must not satisfy this check.
recent_pid=""
for _pid in $(pgrep -f "chromium.*--remote-debugging-port=${CDP_PORT}([[:space:]]|$)" 2>/dev/null); do
# Skip child processes (renderer, gpu, etc.) — only the main browser counts
if ps -o args= -p "$_pid" 2>/dev/null | grep -q -- "--type="; then
continue
fi
_start=$(date -d "$(ps -o lstart= -p "$_pid" 2>/dev/null)" +%s 2>/dev/null || echo 0)
_now=$(date +%s)
if [ $(( _now - _start )) -lt 120 ] && [ "$_start" -gt 0 ]; then
recent_pid="$_pid"
break
fi
done
if [ -n "$recent_pid" ]; then
log "browser launched recently (pid $recent_pid), skipping relaunch (probably still starting)"
exit 0
fi
# Warp-partition recovery (2026-10-04): relaunching Chrome cannot fix a dead
# Warp tunnel — the new browser would fail the same egress check and the
# watchdog would loop. If the failure is warp-egress, restart the tunnel first
# (netvm-node-up.sh is idempotent); only fall through to the Chrome relaunch
# if the tunnel does not recover.
case "$HEALTH_FAIL_REASON" in
"warp egress down"*)
log "warp partition detected ($HEALTH_FAIL_REASON), restarting tunnel via netvm-node-up.sh"
sudo -n "$NETVM_BIN/netvm-node-up.sh" "$PROFILE" >>"$LOG" 2>&1 || true
sleep 5
if healthy; then
log "tunnel restart recovered warp egress, chrome relaunch not needed"
exit 0
fi
log "tunnel restart did not recover egress ($HEALTH_FAIL_REASON), proceeding with chrome relaunch"
;;
esac
log "unhealthy ($HEALTH_FAIL_REASON), relaunching chromebox"
# bracket trick so pkill never matches its own command line
pat="profiles/${PROFILE:0:${#PROFILE}-1}[${PROFILE: -1}]/"
pkill -f "chromium.*$pat" 2>/dev/null || true
sleep 3
# Clean up stale singleton symlinks that break subsequent browser startup
rm -f "/home/super/.local/share/chrome-box/profiles/${PROFILE}/home/.config/chromium/SingletonLock" \
"/home/super/.local/share/chrome-box/profiles/${PROFILE}/home/.config/chromium/SingletonSocket" \
"/home/super/.local/share/chrome-box/profiles/${PROFILE}/home/.config/chromium/SingletonCookie" 2>/dev/null || true
for _sc in $(systemctl --user list-units --type=scope --plain --no-legend 2>/dev/null | awk '{print $1}' | grep -E "^netvm-chrome-${PROFILE}-"); do
systemctl --user stop "$_sc" 2>/dev/null || true
done
# Relaunch in its own systemd scope so it survives this oneshot run.
# nohup/setsid do NOT escape: this timer's service uses KillMode=control-group
# and systemd SIGKILLs everything in the cgroup at teardown (observed
# 2026-10-03: every relaunch "recovered" then died seconds later). A transient
# scope escapes the service cgroup; the scoped process inherits these fds so
# the log redirect below still captures chromium's output.
# NOTE: systemd-run --scope WAITS for the scope's processes (even --no-block,
# verified 2026-10-03), so it must be backgrounded — the scope is an
# independent unit and outlives the wrapper.
# NOTE: --cdp-port is pinned explicitly. netvm-chrome.sh defaults to a
# hash-derived port (9222+...) which will NOT match the registry port that
# muse-chat-api.py uses — a relaunch on the wrong port looks healthy to the
# launcher but is unreachable to the API (observed 2026-10-03: pip relaunched
# on 9278 instead of 9420, watchdog looped on "relaunch FAILED").
# 9>&-: do NOT let the backgrounded launcher inherit the watchdog lock fd.
# Inherited flock fds wedge the lock forever (observed 2026-10-04: opm/muse
# launchers held their profile lock for 100+ min, every later watchdog run
# skipped as "another run in progress" — the watchdog was silently dead).
systemd-run --user --scope --unit="netvm-chrome-${PROFILE}-$(date +%s)" \
"$NETVM_BIN/netvm-chrome.sh" --headless --cdp-port "$CDP_PORT" "$PROFILE" "https://muse.ai" \
>>"$CHROME_LOG" 2>&1 < /dev/null 9>&- &
# Retry with backoff (2026-10-04): cold starts (fresh egress IP, Cloudflare
# handshake) can take >25s for the page title to appear. A single check after
# 25s kills working-but-slow browsers. Try up to 4 times, 15s apart (~60s
# window), logging each attempt. Only declare FAILED if all attempts fail.
_relaunch_ok=0
for _attempt in 1 2 3 4; do
sleep 15
if healthy; then
log "relaunch OK (attempt $_attempt)"
_relaunch_ok=1
break
fi
log "relaunch attempt $_attempt not healthy yet ($HEALTH_FAIL_REASON)"
done
if [ "$_relaunch_ok" -ne 1 ]; then
log "relaunch FAILED ($HEALTH_FAIL_REASON) \u2014 needs operator attention"
exit 1
fi
-312
View File
@@ -1,312 +0,0 @@
#!/usr/bin/env python3
"""Completion auditor: prove work gets done, or say exactly where it stalls.
Runs on a 15-minute systemd timer (systemd/completion-audit.*). Reads
job-log.jsonl, swarms.json, and followups.json; computes the completion
funnel per job family plus swarm drain and followup backlog; writes a JSON
report under logs/ and posts a compact digest to the ops heartbeat sidechat
when degraded (or a heartbeat summary every 6h when green).
Read-only except the digest DM and its own log/state files. Exit 0 always
on a completed audit; tracebacks (real errors) fail the timer visibly.
"""
import argparse
import json
import os
import re
import subprocess
import sys
from collections import Counter, defaultdict
from datetime import datetime, timezone, timedelta
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parent.parent
JOB_LOG = REPO_ROOT / "job-log.jsonl"
SWARMS_FILE = REPO_ROOT / "swarms.json"
FOLLOWUPS_FILE = REPO_ROOT / "followups.json"
LOGS_DIR = REPO_ROOT / "logs"
STATE_FILE = LOGS_DIR / "completion-audit-state.json"
_JOB_ID_RE = re.compile(r"^(.+)-(\d{8})-(\d{6})-([0-9a-f]{8})$")
HEARTBEAT_INTERVAL_H = 6
STALE_RUNNING_MIN = 90
SILENT_MIN_SENT = 3
def utcnow():
return datetime.now(timezone.utc)
def family_of(job_id):
"""Strip the dispatch suffix (<name>-YYYYMMDD-HHMMSS-<hex8>) to the family."""
m = _JOB_ID_RE.match(job_id or "")
return m.group(1) if m else (job_id or "?")
def parse_ts(ts):
try:
t = datetime.fromisoformat(str(ts))
except Exception:
return None
if t.tzinfo is None:
t = t.replace(tzinfo=timezone.utc)
return t
def compute_funnel(events, cutoff):
"""Aggregate job-log events since cutoff.
Returns (families, tools) where families maps family -> counters and
tools holds global tool_exec stats. Pure over the event list.
"""
families = defaultdict(lambda: Counter())
tools = Counter()
tool_errs = Counter()
for e in events:
t = parse_ts(e.get("ts"))
if t is None or t < cutoff:
continue
ty = e.get("type")
if ty == "job_sent":
families[family_of(e.get("job_id"))]["sent"] += 1
elif ty == "job_dispatched":
families[family_of(e.get("job_id"))]["dispatched"] += 1
elif ty == "tool_exec":
tools["total"] += 1
if e.get("success"):
tools["ok"] += 1
else:
tools["fail"] += 1
tool_errs[e.get("op", "?")] += 1
elif ty == "job_result":
fam = family_of(e.get("job_id"))
families[fam]["results"] += 1
families[fam]["ok" if e.get("success") else "fail"] += 1
elif ty == "job_failed":
families[family_of(e.get("job_id"))]["failed"] += 1
elif ty == "fallback_executed":
families[family_of(e.get("job_id"))]["fallback_ok"] += 1
elif ty == "fallback_failed":
families[family_of(e.get("job_id"))]["fallback_fail"] += 1
elif ty == "proof_requested":
families[family_of(e.get("job_id"))]["proofs"] += 1
return families, {"tools": tools, "tool_errs": tool_errs}
def swarm_drain(now):
"""Status counts + stale-running slots from swarms.json."""
try:
data = json.load(open(SWARMS_FILE))
except Exception:
return {"error": "swarms.json unreadable"}, []
values = data.values() if isinstance(data, dict) else data
status = Counter()
stale = []
for s in values:
if not isinstance(s, dict):
continue
for sl in s.get("slots", []) or []:
status[sl.get("status", "?")] += 1
if sl.get("status") == "running":
upd = parse_ts(sl.get("updated_ts"))
if upd and (now - upd) > timedelta(minutes=STALE_RUNNING_MIN):
stale.append({
"swarm": s.get("swarm_id"),
"slot": sl.get("slot"),
"agent": sl.get("agent_id"),
"idle_min": int((now - upd).total_seconds() // 60),
})
return {"slots": dict(status)}, stale
def followup_backlog(now):
"""Pending/overdue/escalated counts from followups.json."""
try:
data = json.load(open(FOLLOWUPS_FILE))
except Exception:
return {"error": "followups.json unreadable"}
values = data.values() if isinstance(data, dict) else data
out = Counter()
for r in values:
if not isinstance(r, dict):
continue
st = r.get("status", "?")
out[st] += 1
if st == "pending":
dl = parse_ts(r.get("deadline"))
if dl and dl < now:
out["overdue"] += 1
return dict(out)
def build_report(hours):
now = utcnow()
cutoff = now - timedelta(hours=hours)
events = []
try:
with open(JOB_LOG) as f:
for line in f:
line = line.strip()
if not line:
continue
try:
events.append(json.loads(line))
except Exception:
continue
except FileNotFoundError:
pass
families, tools = compute_funnel(events, cutoff)
fam = {k: dict(v) for k, v in sorted(families.items())}
swarm, stale = swarm_drain(now)
backlog = followup_backlog(now)
totals = Counter()
for v in fam.values():
for k, n in v.items():
totals[k] += n
silent = sorted(
k for k, v in fam.items()
if v.get("sent", 0) >= SILENT_MIN_SENT and v.get("results", 0) == 0)
degraded_reasons = []
if silent:
degraded_reasons.append(f"{len(silent)} silent families: {', '.join(silent[:5])}")
if totals.get("failed"):
degraded_reasons.append(f"{totals['failed']} job_failed")
if tools["tools"].get("fail"):
top = tools["tool_errs"].most_common(3)
degraded_reasons.append(
"tool errors: " + ", ".join(f"{op}x{n}" for op, n in top))
if totals.get("fallback_fail"):
degraded_reasons.append(f"{totals['fallback_fail']} fallback_failed")
if stale:
degraded_reasons.append(f"{len(stale)} running slots idle >{STALE_RUNNING_MIN}m")
if backlog.get("overdue"):
degraded_reasons.append(f"{backlog['overdue']} overdue followups")
if backlog.get("escalated"):
degraded_reasons.append(f"{backlog['escalated']} escalated followups")
return {
"ts": now.isoformat(),
"window_h": hours,
"totals": dict(totals),
"tools": {k: dict(v) if isinstance(v, Counter) else v
for k, v in tools.items()},
"families": fam,
"silent_families": silent,
"swarms": swarm,
"stale_running": stale[:10],
"followups": backlog,
"degraded": bool(degraded_reasons),
"reasons": degraded_reasons,
}
def render_digest(rep):
t = rep["totals"]
tools = rep["tools"].get("tools", {})
lines = [
f"Completion audit ({rep['window_h']}h, {rep['ts'][:16]}Z)",
f"funnel: {t.get('sent', 0)} sent / {t.get('dispatched', 0)} dispatched / "
f"{tools.get('total', 0)} tool_exec / {t.get('results', 0)} results "
f"({t.get('ok', 0)} ok)",
]
if rep["silent_families"]:
lines.append("silent: " + ", ".join(rep["silent_families"][:6]))
bits = []
if t.get("failed"):
bits.append(f"{t['failed']} job_failed")
if tools.get("fail"):
bits.append(f"{tools['fail']} tool errors")
if t.get("fallback_ok") or t.get("fallback_fail"):
bits.append(f"fallback {t.get('fallback_ok', 0)} ok / {t.get('fallback_fail', 0)} fail")
if t.get("proofs"):
bits.append(f"{t['proofs']} proof reqs")
if bits:
lines.append("flags: " + ", ".join(bits))
sw = rep["swarms"].get("slots", {})
if sw:
lines.append("swarms now: " + " / ".join(f"{v} {k}" for k, v in sorted(sw.items())))
if rep["stale_running"]:
lines.append(f"stale running: {len(rep['stale_running'])} slots (see report)")
fb = rep["followups"]
if fb and "error" not in fb:
lines.append(
f"followups now: {fb.get('pending', 0)} pending / {fb.get('overdue', 0)} "
f"overdue / {fb.get('escalated', 0)} escalated")
if rep["degraded"]:
lines.append("verdict: DEGRADED — " + "; ".join(rep["reasons"][:3]))
else:
lines.append("verdict: HEALTHY — work flowing, results landing")
return "\n".join(lines)
def should_post(report):
"""Post on degraded, else heartbeat at most every HEARTBEAT_INTERVAL_H."""
if report["degraded"]:
return True, "degraded"
try:
state = json.load(open(STATE_FILE))
last = parse_ts(state.get("last_heartbeat"))
except Exception:
last = None
if last is None or (utcnow() - last) > timedelta(hours=HEARTBEAT_INTERVAL_H):
return True, "heartbeat"
return False, "green-quiet"
def post_digest(digest):
argv = [sys.executable, str(REPO_ROOT / "bin" / "dm.py"), "send",
"--agent", "super", "--to", "opm", "--target", "heartbeat", digest]
p = subprocess.run(argv, capture_output=True, text=True, timeout=120)
return p.returncode == 0, (p.stdout or p.stderr or "").strip()[:300]
def save_report(report):
LOGS_DIR.mkdir(parents=True, exist_ok=True)
stamp = report["ts"].replace("+00:00", "Z").replace(":", "")
dated = LOGS_DIR / f"completion-audit-{stamp[:15]}.json"
body = json.dumps(report, indent=2)
dated.write_text(body, encoding="utf-8")
latest = LOGS_DIR / "completion-audit-latest.json"
tmp = LOGS_DIR / f".completion-audit-latest.tmp.{os.getpid()}"
tmp.write_text(body, encoding="utf-8")
os.replace(tmp, latest)
return dated
def main():
ap = argparse.ArgumentParser(description="Completion funnel auditor")
ap.add_argument("--hours", type=int, default=24)
ap.add_argument("--post", dest="post", action="store_true", default=True)
ap.add_argument("--no-post", dest="post", action="store_false")
ap.add_argument("--json", action="store_true", help="Print raw report JSON")
args = ap.parse_args()
report = build_report(args.hours)
path = save_report(report)
if args.json:
print(json.dumps(report, indent=2))
else:
print(render_digest(report))
print(f"\nreport: {path}")
if not args.post:
print("post: skipped (--no-post)")
return 0
do_post, why = should_post(report)
if not do_post:
print(f"post: skipped ({why})")
return 0
ok, detail = post_digest(render_digest(report))
print(f"post: {'delivered' if ok else 'FAILED'} ({why}) {detail[:120]}")
if ok and why == "heartbeat":
try:
state = {}
if STATE_FILE.exists():
state = json.loads(STATE_FILE.read_text(encoding="utf-8"))
state["last_heartbeat"] = report["ts"]
STATE_FILE.write_text(json.dumps(state, indent=2), encoding="utf-8")
except Exception as e:
print(f"warning: state save failed: {e}")
return 0
if __name__ == "__main__":
sys.exit(main())
-416
View File
@@ -1,416 +0,0 @@
#!/usr/bin/env python3
"""cred-client.py — Agent API client & CLI for cred.muse-dev.online.
Provides programmatic and CLI access for agents and operators to:
- initiate: start onboarding for a client email/node
- submit-otp: submit verification code transiently
- status: query onboarding & vitality state for a node
- list: list fleet accounts, emails, and statuses
Authentication:
- Machine identity signature via ~/.ssh/muse-health (machine bl)
- Or Bearer token from CRED_TOKEN / OPERATOR_TOKEN environment variable.
Usage:
cred-client.py initiate --node <node> --email <email> [--service muse] [--account-name <name>]
cred-client.py submit-otp --node <node> --otp <code> [--email <email>]
cred-client.py status --node <node> [--json]
cred-client.py list [--json]
"""
import argparse
import datetime
import importlib.util
import json
import os
import re
import subprocess
import sys
import tempfile
import urllib.error
import urllib.request
NETVM_DIR = os.environ.get("NETVM_DIR", "/home/super/Projects/NetVM")
KEY_DEFAULT = os.path.expanduser("~/.ssh/muse-health")
CRED_API_DEFAULT = os.environ.get("CRED_API_URL", "https://cred.muse-dev.online/api/cred")
def load_registry():
path = os.path.join(NETVM_DIR, "bin", "netvm-registry.py")
try:
spec = importlib.util.spec_from_file_location("netvm_registry", path)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return mod
except Exception:
return None
def get_accounts():
"""Parse ACCOUNTS.md into a structured list of accounts."""
path = os.path.join(NETVM_DIR, "ACCOUNTS.md")
accounts = []
if not os.path.exists(path):
return accounts
with open(path) as f:
for line in f:
line = line.strip()
if not line.startswith("|") or line.startswith("| agent") or line.startswith("|-------"):
continue
cols = [c.strip() for c in line.strip("|").split("|")]
if len(cols) >= 9:
accounts.append({
"agent": cols[0],
"node": cols[1],
"profile": cols[2],
"login_type": cols[3],
"meta_label": cols[4],
"email": cols[5],
"phone_otp": cols[6],
"instagram_linked": cols[7],
"status": cols[8],
"egress_ip": cols[9] if len(cols) > 9 else "",
"cdp_port": cols[10] if len(cols) > 10 else "",
"display_name": cols[11] if len(cols) > 11 else "",
"notes": cols[12] if len(cols) > 12 else ""
})
return accounts
def sign_payload(payload_bytes, key_path=KEY_DEFAULT, namespace="cred"):
"""Sign payload using SSH ed25519 key (zero secret across wire)."""
if not os.path.exists(key_path):
return None
tmp = tempfile.mkdtemp()
try:
data_file = os.path.join(tmp, "data")
with open(data_file, "wb") as f:
f.write(payload_bytes)
r = subprocess.run(
["ssh-keygen", "-Y", "sign", "-f", key_path, "-n", namespace, data_file],
capture_output=True, text=True
)
if r.returncode != 0:
return None
with open(data_file + ".sig") as f:
return f.read().strip()
finally:
subprocess.run(["rm", "-rf", tmp], capture_output=True)
class CredClient:
def __init__(self, api_url=CRED_API_DEFAULT, key_path=KEY_DEFAULT):
self.api_url = api_url.rstrip("/")
self.key_path = key_path
self.token = os.environ.get("CRED_TOKEN") or os.environ.get("OPERATOR_TOKEN")
def run_local_driver(self, node, step, email, otp=None, account_name=None):
"""Execute local onboard-driver.py inside the node's netns."""
exec_script = os.path.join(NETVM_DIR, "bin", "netvm-exec.sh")
driver_script = os.path.join(NETVM_DIR, "bin", "onboard-driver.py")
cmd = [exec_script, node, "--", sys.executable, driver_script,
"--node", node, "--service", "muse", "--id-type", "email", "--step", step]
if account_name:
cmd += ["--account-name", account_name]
input_data = email + "\n"
if otp:
input_data += str(otp) + "\n"
p = subprocess.run(cmd, input=input_data, capture_output=True, text=True, timeout=240)
return p.returncode, p.stdout.strip(), p.stderr.strip()
def initiate(self, node, email, service="muse", account_name=None):
"""Initiate client onboarding."""
ret, stdout, stderr = self.run_local_driver(node, "initiate", email, account_name=account_name)
if ret == 0:
return {"status": "active", "node": node, "email": email, "message": "Already authenticated and session active."}
elif ret == 2:
return {"status": "awaiting_otp", "node": node, "email": email, "message": "OTP verification code sent. Awaiting input."}
elif ret == 3:
return {"status": "needs_human", "node": node, "email": email, "message": "Multiple accounts match identifier. Manual selection required."}
else:
return {"status": "error", "node": node, "email": email, "code": ret, "detail": stderr or stdout}
def submit_otp(self, node, otp, email=None, service="muse"):
"""Submit transient verification code."""
if not email:
# Look up email from ACCOUNTS.md
for acct in get_accounts():
if acct["node"] == node and acct["email"] not in ("-", ""):
email = acct["email"]
break
if not email:
email = "unknown"
ret, stdout, stderr = self.run_local_driver(node, "submit", email, otp=otp)
if ret == 0:
return {"status": "active", "node": node, "email": email, "message": "Authentication successful. Chat session active."}
elif ret == 3:
return {"status": "needs_human", "node": node, "email": email, "message": "Multiple accounts match identifier."}
else:
return {"status": "error", "node": node, "email": email, "code": ret, "detail": stderr or stdout}
def status(self, node):
"""Check status of a node."""
accounts = get_accounts()
target = next((a for a in accounts if a["node"] == node), None)
port = None
reg = load_registry()
if reg:
port = reg.port_for(node)
# Vitality check
alive = False
detail = ""
if port:
try:
checker = os.path.join(NETVM_DIR, "bin", "accounts-health.py")
cmd = ["sudo", "-n", "ip", "netns", "exec", f"warp-{node}", sys.executable, checker, str(port)]
r = subprocess.run(cmd, capture_output=True, text=True, timeout=10)
if r.returncode == 0:
data = json.loads(r.stdout.strip().splitlines()[-1])
alive = data.get("session_alive", False)
detail = data.get("detail", "")
except Exception as e:
detail = str(e)
return {
"node": node,
"status": target["status"] if target else "unregistered",
"email": target["email"] if target else "-",
"cdp_port": port,
"session_alive": alive,
"detail": detail
}
def list_all(self):
"""List all accounts and live statuses."""
res = []
for a in get_accounts():
st = self.status(a["node"])
res.append(st)
return res
def notify_operator(self, subject, body, to_email="defnotabotnet@gmail.com"):
"""Send notification to operator via local MTA (msmtp or mail)."""
sent = False
err = None
# Try msmtp first
try:
p = subprocess.Popen(["msmtp", to_email], stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True)
msg = f"Subject: {subject}\nTo: {to_email}\nFrom: NetVM Automation <root@bl>\n\n{body}\n"
stdout, stderr = p.communicate(input=msg)
if p.returncode == 0:
sent = True
else:
err = stderr.strip() or stdout.strip()
except Exception as e:
err = str(e)
if not sent:
# Fallback to mail / s-nail
try:
p = subprocess.Popen(["mail", "-s", subject, to_email], stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True)
stdout, stderr = p.communicate(input=body)
if p.returncode == 0:
sent = True
else:
err = stderr.strip() or stdout.strip()
except Exception as e:
err = str(e)
return {"sent": sent, "recipient": to_email, "error": err if not sent else None}
def link_instagram(self, node, notify=False):
"""Fetch the Meta Accounts Center OAuth URL and provide one-tap Tailscale link."""
portal_script = os.path.join(NETVM_DIR, "bin", "tailscale-verify-portal.py")
ts_ip = "100.123.153.75"
ts_dns = "bl.tailfb5960.ts.net"
try:
import importlib.util
spec = importlib.util.spec_from_file_location("portal", portal_script)
portal_mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(portal_mod)
ts_ip = portal_mod.get_tailscale_ip()
ts_dns = portal_mod.get_tailscale_dns()
ig_data = portal_mod.fetch_node_ig_link(node)
except Exception as e:
ig_data = {"error": str(e)}
direct_url = ig_data.get("url") if isinstance(ig_data, dict) else None
portal_url = f"http://{ts_dns}:8765/verify/{node}"
portal_ip_url = f"http://{ts_ip}:8765/verify/{node}"
email_result = None
if notify and direct_url:
subject = f"[NetVM Action Required] Link Instagram for Node '{node}'"
body = (
f"Node '{node}' is waiting at the Muse age verification gate.\n\n"
f"Tap the Tailscale portal link from your device to approve:\n"
f" {portal_url}\n"
f" (or {portal_ip_url})\n\n"
f"Direct OAuth URL:\n"
f" {direct_url}\n\n"
f"After approving in Instagram, NetVM will automatically transition node '{node}' to active."
)
email_result = self.notify_operator(subject, body)
return {
"node": node,
"status": "needs_verification",
"portal_url": portal_url,
"portal_ip_url": portal_ip_url,
"direct_oauth_url": direct_url,
"error": ig_data.get("error") if isinstance(ig_data, dict) else None,
"email_notified": email_result.get("sent") if email_result else False
}
def audit_meta(self, node):
"""Query Meta Accounts Center for linked profiles and security status."""
meta_script = os.path.join(NETVM_DIR, "bin", "meta-acct.py")
cmd = ["sudo", "-n", "ip", "netns", "exec", f"warp-{node}", sys.executable, meta_script, "list-linked", node]
try:
res = subprocess.run(cmd, capture_output=True, text=True, timeout=15)
out = res.stdout.strip()
if out:
# Find outermost JSON object
start = out.find("{")
end = out.rfind("}")
if start >= 0 and end >= start:
return json.loads(out[start:end+1])
return json.loads(out)
return {"error": res.stderr.strip() or "No output from meta-acct.py"}
except Exception as e:
return {"error": str(e)}
def main():
common = argparse.ArgumentParser(add_help=False)
common.add_argument("--json", action="store_true", help="Output JSON response")
p = argparse.ArgumentParser(description="cred-client: Agent API client for cred onboarding", parents=[common])
sub = p.add_subparsers(dest="command", required=True)
# initiate
p_init = sub.add_parser("initiate", parents=[common], help="Initiate onboarding flow for a node")
p_init.add_argument("--node", required=True, help="Node label (e.g. opm, pip, client1)")
p_init.add_argument("--email", required=True, help="Client login email")
p_init.add_argument("--service", default="muse", help="Service name (default: muse)")
p_init.add_argument("--account-name", default=None, help="Display name hint for multi-account selector")
# submit-otp
p_otp = sub.add_parser("submit-otp", parents=[common], help="Submit transient OTP verification code")
p_otp.add_argument("--node", required=True, help="Node label")
p_otp.add_argument("--otp", required=True, help="6-digit verification code")
p_otp.add_argument("--email", default=None, help="Client email (optional, auto-detected from registry)")
# status
p_stat = sub.add_parser("status", parents=[common], help="Query onboarding and session status for a node")
p_stat.add_argument("--node", required=True, help="Node label")
# link-instagram
p_link = sub.add_parser("link-instagram", parents=[common], help="Generate Tailscale one-tap portal & OAuth link for age verification")
p_link.add_argument("--node", required=True, help="Node label")
p_link.add_argument("--notify", action="store_true", help="Send email alert to operator via local MTA")
# meta-audit
p_meta = sub.add_parser("meta-audit", parents=[common], help="Query Meta Accounts Center for linked profiles and security status")
p_meta.add_argument("--node", required=True, help="Node label")
# list
sub.add_parser("list", parents=[common], help="List all registered nodes and vitality statuses")
args = p.parse_args()
client = CredClient()
if args.command == "meta-audit":
res = client.audit_meta(args.node)
if args.json:
print(json.dumps(res, indent=2))
else:
print(f"\n=== META ACCOUNTS CENTER AUDIT: {args.node} ===")
if res.get("error"):
print(f"Error: {res['error']}", file=sys.stderr)
sys.exit(1)
print(f"Meta Account Email: {res.get('email') or '(none / phone-only)'}")
profiles = res.get("profiles", [])
print(f"Linked Profiles ({len(profiles)}):")
for p_info in profiles:
print(f" - [{p_info.get('type')}] {p_info.get('name')}")
sys.exit(0)
if args.command == "link-instagram":
res = client.link_instagram(args.node, notify=args.notify)
if args.json:
print(json.dumps(res, indent=2))
else:
print(f"\n=== INSTAGRAM LINKING: {args.node} ===")
if res.get("error"):
print(f"Error: {res['error']}", file=sys.stderr)
sys.exit(1)
print(f"Tailscale Portal: {res['portal_url']}")
print(f"Direct IP Portal: {res['portal_ip_url']}")
print(f"OAuth URL: {res['direct_oauth_url'][:80]}...")
if args.notify:
print(f"Email Notified: {'YES' if res['email_notified'] else 'FAILED'}")
print("\nTap either Tailscale link from your phone/browser to complete Meta age verification.")
sys.exit(0)
if args.command == "initiate":
res = client.initiate(args.node, args.email, service=args.service, account_name=args.account_name)
if args.json:
print(json.dumps(res, indent=2))
else:
st = res.get("status")
if st == "awaiting_otp":
print(f"[APPROVAL_NEEDED] Code sent to [redacted] for node '{args.node}'.")
print(f"Submit OTP with: super cred submit-otp --node {args.node} --otp <code>")
sys.exit(2)
elif st == "active":
print(f"[SUCCESS] Node '{args.node}' is already authenticated and active.")
sys.exit(0)
else:
print(f"[{st.upper()}] {res.get('message') or res.get('detail')}")
sys.exit(res.get("code", 1))
elif args.command == "submit-otp":
res = client.submit_otp(args.node, args.otp, email=args.email)
if args.json:
print(json.dumps(res, indent=2))
else:
st = res.get("status")
if st == "active":
print(f"[SUCCESS] Node '{args.node}' successfully signed in!")
sys.exit(0)
else:
print(f"[{st.upper()}] {res.get('message') or res.get('detail')}")
sys.exit(res.get("code", 1))
elif args.command == "status":
res = client.status(args.node)
if args.json:
print(json.dumps(res, indent=2))
else:
print(f"Node: {res['node']}")
print(f"Email: {res['email']}")
print(f"Status: {res['status']}")
print(f"CDP Port: {res['cdp_port'] or '-'}")
print(f"Session Alive: {'YES' if res['session_alive'] else 'NO'}")
if res["detail"]:
print(f"Detail: {res['detail']}")
elif args.command == "list":
items = client.list_all()
if args.json:
print(json.dumps(items, indent=2))
else:
print(f"{'NODE':<10} {'STATUS':<14} {'CDP':<6} {'ALIVE':<7} {'EMAIL'}")
print("-" * 65)
for it in items:
alive_str = "YES" if it["session_alive"] else "NO"
print(f"{it['node']:<10} {it['status']:<14} {str(it['cdp_port'] or '-'):<6} {alive_str:<7} {it['email']}")
if __name__ == "__main__":
main()
-1
View File
@@ -1 +0,0 @@
cred-client.py
-178
View File
@@ -1,178 +0,0 @@
#!/usr/bin/env python3
"""
crypt-server.py — Public cryptographic attestation and key directory for NetVM.
Serves https://crypt.muse-dev.online/ (via cloudflared / reverse proxy).
Endpoints:
GET / -> Service directory / health JSON
GET /health -> Health check
GET /keys/allowed_signers -> OpenSSH allowed_signers formatted file
GET /keys/{identity}.pub -> Individual public key
GET /proofs -> List known proof hashes / work order attestations
GET /proofs/{id} -> Retrieve proof envelope and SSH signature
POST /proofs -> Submit / register a signed proof record
"""
import argparse
import glob
import json
import os
import re
import ssl
import sys
from http.server import HTTPServer, BaseHTTPRequestHandler
REPO_DIR = "/home/super/Projects/NetVM"
SIGNERS_DIR = os.path.join(REPO_DIR, "dm-signers")
PROOFS_DIR = os.path.join(REPO_DIR, "var", "proofs")
os.makedirs(PROOFS_DIR, exist_ok=True)
class CryptHandler(BaseHTTPRequestHandler):
server_version = "crypt-attestation/1.0"
def log_message(self, format, *args):
sys.stderr.write(f"crypt-server: {self.client_address[0]} - {format % args}\n")
def _send(self, code, content, content_type="application/json"):
if isinstance(content, (dict, list)):
body = json.dumps(content, indent=2).encode("utf-8")
elif isinstance(content, str):
body = content.encode("utf-8")
else:
body = bytes(content)
self.send_response(code)
self.send_header("Content-Type", content_type)
self.send_header("Content-Length", str(len(body)))
self.send_header("Access-Control-Allow-Origin", "*")
self.end_headers()
self.wfile.write(body)
def do_GET(self):
path = self.path.split("?")[0].rstrip("/")
if not path:
path = "/"
if path in ("/", "/health"):
self._send(200, {
"service": "crypt.muse-dev.online",
"status": "active",
"mode": "public_attestation",
"endpoints": [
"/keys/allowed_signers",
"/keys/<identity>.pub",
"/proofs",
"/proofs/<id>"
]
})
return
if path == "/keys/allowed_signers":
allowed_path = os.path.join(SIGNERS_DIR, "allowed_signers")
if os.path.exists(allowed_path):
with open(allowed_path, "r") as f:
data = f.read()
self._send(200, data, content_type="text/plain; charset=utf-8")
else:
self._send(404, {"error": "allowed_signers not found"})
return
m_key = re.match(r"^/keys/([a-zA-Z0-9_\-\.]+)\.pub$", path)
if m_key:
ident = m_key.group(1)
pub_path = os.path.join(SIGNERS_DIR, f"{ident}.pub")
if os.path.exists(pub_path):
with open(pub_path, "r") as f:
data = f.read()
self._send(200, data, content_type="text/plain; charset=utf-8")
else:
self._send(404, {"error": f"Public key for {ident} not found"})
return
if path == "/proofs":
proof_files = glob.glob(os.path.join(PROOFS_DIR, "*.json"))
ids = [os.path.basename(p)[:-5] for p in proof_files]
self._send(200, {"proofs": sorted(ids)})
return
m_proof = re.match(r"^/proofs/([a-zA-Z0-9_\-]+)$", path)
if m_proof:
pid = m_proof.group(1)
pf = os.path.join(PROOFS_DIR, f"{pid}.json")
if os.path.exists(pf):
with open(pf, "r") as f:
data = json.load(f)
self._send(200, data)
else:
self._send(404, {"error": f"Proof {pid} not found"})
return
self._send(404, {"error": "not found"})
def do_POST(self):
path = self.path.split("?")[0].rstrip("/")
if path == "/proofs":
try:
length = int(self.headers.get("Content-Length", 0))
except ValueError:
length = 0
if length <= 0 or length > 65536:
self._send(400, {"error": "invalid content length"})
return
try:
data = json.loads(self.rfile.read(length))
except Exception:
self._send(400, {"error": "malformed JSON"})
return
pid = data.get("id")
if not pid or not re.match(r"^[a-zA-Z0-9_\-]+$", pid):
self._send(400, {"error": "missing or invalid proof id"})
return
pf = os.path.join(PROOFS_DIR, f"{pid}.json")
with open(pf, "w") as f:
json.dump(data, f, indent=2)
self._send(201, {"status": "stored", "id": pid, "url": f"https://crypt.muse-dev.online/proofs/{pid}"})
return
self._send(404, {"error": "not found"})
def main():
parser = argparse.ArgumentParser(description="Crypt attestation and key server")
parser.add_argument("--port", type=int, default=8446)
parser.add_argument("--host", default="100.123.153.75")
parser.add_argument("--ssl", action="store_true", help="Enable self-signed HTTPS")
args = parser.parse_args()
server = HTTPServer((args.host, args.port), CryptHandler)
if args.ssl:
cert_file = os.path.join(REPO_DIR, "ssl", "crypt-selfsigned.crt")
key_file = os.path.join(REPO_DIR, "ssl", "crypt-selfsigned.key")
os.makedirs(os.path.dirname(cert_file), exist_ok=True)
if not os.path.exists(cert_file):
import subprocess
subprocess.run([
"openssl", "req", "-x509", "-newkey", "rsa:2048",
"-keyout", key_file, "-out", cert_file,
"-days", "3650", "-nodes",
"-subj", "/CN=crypt.muse-dev.online"
], check=True, capture_output=True)
os.chmod(key_file, 0o600)
ctx = ssl.SSLContext(ssl.PROTOCOL_TLS_SERVER)
ctx.load_cert_chain(cert_file, key_file)
server.socket = ctx.wrap_socket(server.socket, server_side=True)
print(f"Crypt server running on {'https' if args.ssl else 'http'}://{args.host}:{args.port}", file=sys.stderr)
try:
server.serve_forever()
except KeyboardInterrupt:
print("\nShutting down crypt server", file=sys.stderr)
if __name__ == "__main__":
main()
-266
View File
@@ -1,266 +0,0 @@
#!/usr/bin/env python3
"""
Side-chat to main-chat work siphon — detection rules (REPAIRED, agent 2 of 5).
Fixes the false-positive ✅ COMPLETED relay at the source:
"Sending is disabled until this conversation can be verified." → COMPLETED
was caused by a SINGLE keyword ("verified") matching one regex.
Repairs (see OUTPUT.md for rationale):
1. COMPLETED requires >= 2 DISTINCT pattern hits (weighted: the structured
`[RESULT ...] OK` marker counts 2 — it is the fleet's own machine-emitted
completion signal, far less ambiguous than a bare "done").
2. Negation guards: negation/failure-state words veto COMPLETED outright
(fail-closed: a negated completion claim is never relayed as complete).
3. Honest labeling: the fake "confidence 60%" (which literally meant "one
regex hit") is replaced by a keyword-hit count. SiphonHit.hits is the
authoritative field; `confidence` is kept for backward compatibility
but must NOT be rendered as a percentage anywhere user-facing.
4. Stale suppression: a message older than 15 minutes never relays as
COMPLETED. Pass message_ts (epoch seconds). monitor.py currently does
NOT pass a timestamp — agent 3 / the integrator must thread
message["ts"] through (see OUTPUT.md).
DO NOT overwrite the original detect.py with this file until the integrator
reconciles all 5 agents' outputs.
"""
import re
import time
from dataclasses import dataclass, field
from typing import Optional
@dataclass
class SiphonHit:
category: str # COMPLETED, BLOCKER, DECISION, ALERT, MILESTONE
confidence: float # LEGACY — kept for API compatibility only.
# Do NOT render as "confidence NN%"; it is not a
# reliability measure. See `hits`.
hits: int = 0 # AUTHORITATIVE — distinct keyword-pattern hits
# (weighted; see COMPLETED_PATTERN_WEIGHTS).
summary: str = "" # one-line summary for main chat
thread_id: str = "" # source side chat
message_id: str = "" # source message
author: str = "" # INTEGRATOR (agent 3 absent) — the message's real
# author, plumbed from message["author"] by
# monitor.py. Empty = unknown; NEVER substitute the
# thread's registered agent silently (see
# format_siphon).
# Full message text is NOT stored here — main chat gets a summary
# plus a link back, never the full content (safety: no sensitive
# data siphoned verbatim).
# --- Keyword sets (configurable) ---
COMPLETED_PATTERNS = [
re.compile(r'\b(done|completed|finished|deployed|shipped|live|verified)\b', re.I),
re.compile(r'\b(all|tests?)\s+(pass|green|passing)\b', re.I),
re.compile(r'\[RESULT[^\]]*\]\s*OK', re.I),
re.compile(r'\b(merged|committed|pushed|published)\b', re.I),
]
# Weighted hits: the structured [RESULT] OK marker is the fleet's own
# machine-emitted completion signal — unambiguous enough to stand alone.
COMPLETED_PATTERN_WEIGHTS = {0: 1, 1: 1, 2: 2, 3: 1}
COMPLETED_MIN_WEIGHT = 2 # >= 2 distinct pattern hits (or one [RESULT] OK)
BLOCKER_PATTERNS = [
re.compile(r'\b(blocked|stuck|failing|broken|down|error|failed)\b', re.I),
re.compile(r'\b(need|needs|waiting)\s+(your|approval|input|decision)\b', re.I),
re.compile(r'\b(can\'t|cannot|unable to)\b', re.I),
re.compile(r'\[RESULT[^\]]*\]\s*(FAIL|ERROR)', re.I),
]
DECISION_PATTERNS = [
re.compile(r'\b(should (i|we)|shall i|want me to)\b', re.I),
re.compile(r'\b(your call|needs? your|awaiting your)\b', re.I),
re.compile(r'\b(approve|approval)\b.*\?', re.I),
re.compile(r'^(yes|no)\s*\?\s*$', re.I),
]
ALERT_PATTERNS = [
re.compile(r'\b(security|vulnerability|breach|compromised|exploit)\b', re.I),
re.compile(r'\b(urgent|critical|emergency|asap)\b', re.I),
re.compile(r'\b(502|503|500)\b.*\b(error|down)\b', re.I),
re.compile(r'\b(ssh|tunnel).*\b(down|broken|failed)\b', re.I),
]
MILESTONE_PATTERNS = [
re.compile(r'\b(milestone|phase \d+ (complete|done)|shipped v)\b', re.I),
re.compile(r'\b(all \d+ (items? )?done)\b', re.I),
]
# Patterns that suppress siphoning (safety)
SUPPRESS_PATTERNS = [
re.compile(r'\b(password|secret|token|key|pin)\s*[:=]', re.I),
re.compile(r'-----BEGIN', re.I), # never siphon key material
re.compile(r'\[do not siphon\]', re.I), # explicit opt-out marker
]
# --- Negation guards: any match vetoes COMPLETED (fail-closed) ---
# A completion claim in the presence of negation / failure-state language
# is never relayed as ✅ COMPLETED, no matter how many keywords hit.
NEGATION_GUARDS = [
# explicit negation
re.compile(r'\b(not|never|no|nothing|none|neither|nor)\b', re.I),
re.compile(r"\b(do not|don't|didn't|doesn't|won't|can't|cannot|isn't|aren't|"
r"wasn't|weren't|haven't|hasn't|hadn't|couldn't|shouldn't)\b", re.I),
# incompleteness hedges
re.compile(r'\b(still|yet|pending|unfinished|incomplete)\b', re.I),
# failure-state words (a "completed" message containing these is suspect)
re.compile(r'\b(broken|failed|failing|failure|down|stuck|blocked|disabled|'
r'error|errors|crash|crashed)\b', re.I),
# hedging conjunctions ("deployed, but tests are red")
re.compile(r'\b(but|however|although|though)\b', re.I),
]
# Messages older than this never relay as COMPLETED (seconds).
COMPLETED_MAX_AGE_S = 15 * 60
def _match_score(text: str, patterns) -> float:
"""Legacy confidence for non-COMPLETED categories (unchanged)."""
hits = sum(1 for p in patterns if p.search(text))
if hits == 0:
return 0.0
# Diminishing returns: 1 hit = 0.6, 2 = 0.8, 3+ = 0.95
return min(0.95, 0.6 + (hits - 1) * 0.2)
def _completed_weight(text: str):
"""
Return (weighted_hits, distinct_hits, matched_pattern_indexes) for
COMPLETED_PATTERNS. Weighted: [RESULT] OK counts 2.
"""
matched = [i for i, p in enumerate(COMPLETED_PATTERNS) if p.search(text)]
weight = sum(COMPLETED_PATTERN_WEIGHTS.get(i, 1) for i in matched)
return weight, len(matched), matched
def _is_negated(text: str) -> bool:
"""True if any negation guard fires anywhere in the text."""
return any(p.search(text) for p in NEGATION_GUARDS)
def _extract_summary(text: str, max_len: int = 120) -> str:
"""Extract a safe one-line summary. Strips to first meaningful line."""
# Take first non-empty line, truncate
for line in text.strip().split('\n'):
line = line.strip()
if line and len(line) > 10:
if len(line) > max_len:
return line[:max_len - 3] + '...'
return line
return text[:max_len]
def detect(text: str, thread_id: str, message_id: str,
min_confidence: float = 0.6,
message_ts: Optional[float] = None) -> Optional[SiphonHit]:
"""
Check a side chat message for siphon-worthy content.
Returns SiphonHit or None.
message_ts: epoch seconds of the original message (optional). Messages
older than COMPLETED_MAX_AGE_S (15 min) never relay as COMPLETED.
NOTE: monitor.py does not currently pass a timestamp — agent 3 / the
integrator must thread message["ts"] through the detect() call.
"""
# Safety: suppress sensitive content
for p in SUPPRESS_PATTERNS:
if p.search(text):
return None
# Stale suppression applies to COMPLETED only.
completed_allowed = True
if message_ts is not None:
try:
age = time.time() - float(message_ts)
if age > COMPLETED_MAX_AGE_S:
completed_allowed = False
except (TypeError, ValueError):
pass # unparseable ts: proceed, do not fail closed on metadata
# COMPLETED: >=2 distinct weighted pattern hits, no negation, not stale.
completed_hits = 0
completed_conf = 0.0
if completed_allowed and not _is_negated(text):
weight, distinct, _ = _completed_weight(text)
if weight >= COMPLETED_MIN_WEIGHT:
completed_hits = weight
# legacy confidence kept for API compat; NOT a reliability measure
completed_conf = min(0.95, 0.6 + (distinct - 1) * 0.2)
candidates = [
("COMPLETED", completed_conf, completed_hits),
("BLOCKER", _match_score(text, BLOCKER_PATTERNS),
sum(1 for p in BLOCKER_PATTERNS if p.search(text))),
("DECISION", _match_score(text, DECISION_PATTERNS),
sum(1 for p in DECISION_PATTERNS if p.search(text))),
("ALERT", _match_score(text, ALERT_PATTERNS),
sum(1 for p in ALERT_PATTERNS if p.search(text))),
("MILESTONE", _match_score(text, MILESTONE_PATTERNS),
sum(1 for p in MILESTONE_PATTERNS if p.search(text))),
]
# Sort by confidence descending; ALERT wins ties (safety: urgency first)
# Use negative confidence for descending, and ALERT as tiebreaker
candidates.sort(key=lambda x: (-x[1], 0 if x[0] == "ALERT" else 1))
best_cat, best_conf, best_hits = candidates[0]
if best_conf < min_confidence:
return None
return SiphonHit(
category=best_cat,
confidence=best_conf,
hits=best_hits,
summary=_extract_summary(text),
thread_id=thread_id,
message_id=message_id,
)
def format_siphon(hit: SiphonHit, agent_name: str = "sidechat") -> str:
"""
Format a siphon message for main chat.
HONEST LABELING: reports keyword hit count, never a fake "confidence %".
HONEST AUTHORSHIP (integrator, agent 3 absent): attributes the message's
real author when known. Falls back to the thread's registered agent only
when the author is unknown — and says so explicitly, so a relay can
never again launder thread ownership as authorship.
"""
emoji = {"COMPLETED": "✅", "BLOCKER": "🚧", "DECISION": "❓",
"ALERT": "🚨", "MILESTONE": "🎯"}.get(hit.category, "📋")
thread_url = f"https://muse.ai/thread/{hit.thread_id}"
if hit.author:
attribution = f"from {hit.author}"
else:
attribution = f"from {agent_name} side chat (author unverified)"
return (
f"{emoji} [{hit.category}] {attribution}\n"
f"{hit.summary}\n"
f"→ {thread_url}\n"
f"(keyword hits: {hit.hits})"
)
# --- Opt-out registry ---
_opt_out_threads: set = set()
def opt_out(thread_id: str):
"""Agent opts a side chat out of siphoning."""
_opt_out_threads.add(thread_id)
def opt_in(thread_id: str):
"""Re-enable siphoning for a side chat."""
_opt_out_threads.discard(thread_id)
def is_opted_out(thread_id: str) -> bool:
return thread_id in _opt_out_threads
-74
View File
@@ -1,74 +0,0 @@
#!/usr/bin/env python3
"""
Side-chat work digest — INTEGRATOR minimal version (agent 4 of 5 never
delivered its digest/throttle design within the window).
Purpose: stop the per-message ✅ COMPLETED relay flood. COMPLETED hits are
batched here and emitted as ONE periodic digest instead of N main-chat
messages. ALERT / BLOCKER / DECISION still relay individually via
siphon() — urgency is never batched.
Interface (what agent 4's full design should remain compatible with):
- buffer = DigestBuffer(max_items=20, max_age_s=3600)
- buffer.add(hit) -> None
- buffer.flush() -> Optional[str] (formatted digest, clears buffer)
- flush_digest() -> Optional[str] (module-level singleton convenience)
Reversible: to restore per-message COMPLETED relays, route COMPLETED back
through siphon() in monitor.py and ignore this module.
"""
import time
from typing import List, Optional
try:
from detect import SiphonHit
except ImportError: # pragma: no cover
SiphonHit = object
class DigestBuffer:
"""Batch COMPLETED hits; flush() renders one digest message."""
def __init__(self, max_items: int = 20, max_age_s: int = 3600):
self.max_items = max_items
self.max_age_s = max_age_s
self._items: List[tuple] = [] # (ts, SiphonHit)
def add(self, hit) -> None:
now = time.time()
# Prune items older than max_age_s on every add (bounded memory).
self._items = [(ts, h) for ts, h in self._items
if now - ts < self.max_age_s]
self._items.append((now, hit))
# Bound the buffer; oldest evicted first.
self._items = self._items[-self.max_items:]
def __len__(self) -> int:
return len(self._items)
def flush(self) -> Optional[str]:
"""Render and clear. Returns None when there's nothing to digest."""
if not self._items:
return None
lines = ["📦 [DIGEST] completions from side chats "
f"({len(self._items)} item(s))"]
for _, hit in self._items:
author = getattr(hit, "author", "") or "?"
url = f"https://muse.ai/thread/{hit.thread_id}"
lines.append(f"• {hit.summary} — {author} ({url})")
self._items = []
return "\n".join(lines)
# Module-level singleton: the monitor loop shares one buffer per process.
_default_buffer = DigestBuffer()
def get_buffer() -> DigestBuffer:
return _default_buffer
def flush_digest() -> Optional[str]:
"""Flush the process-wide digest buffer. None if empty."""
return _default_buffer.flush()
-267
View File
@@ -1,267 +0,0 @@
#!/usr/bin/env python3
"""dm-log-taxonomy.py — READ-ONLY failure taxonomy for fleet DM sidechat reliability.
Reads /home/super/Projects/NetVM/dm-log.jsonl, prints:
1. Event-type counts and send outcome rates
2. Sidechat nav failure taxonomy (per-target, per-agent-pair)
3. Failure timeline (hourly buckets, worst 10-min windows, by node)
4. "Ghost" rate: verified:true sidechat sends with no UUID anywhere
5. Hypothesis evidence tables (nav_failed reasons, placement pairs,
alias sources, autoprovision success, retry distribution)
No writes to any state files. Runs in <1s on the current log size.
"""
import json
import re
import sys
from collections import Counter, defaultdict
from datetime import datetime, timezone, timedelta
LOG = "/home/super/Projects/NetVM/dm-log.jsonl"
WINDOW_H = 48
def parse_ts(s):
if not s:
return None
try:
dt = datetime.fromisoformat(s.replace("Z", "+00:00"))
if dt.tzinfo is None:
dt = dt.replace(tzinfo=timezone.utc)
return dt
except Exception:
return None
def main():
now = datetime.now(timezone.utc)
cutoff = now - timedelta(hours=WINDOW_H)
events = []
parse_err = 0
with open(LOG, encoding="utf-8") as f:
for line in f:
line = line.strip()
if not line:
continue
try:
e = json.loads(line)
except Exception:
parse_err += 1
continue
e["_dt"] = parse_ts(e.get("ts"))
events.append(e)
in_win = [e for e in events if e["_dt"] and e["_dt"] >= cutoff]
print(f"dm-log.jsonl: {len(events)} total lines ({parse_err} parse errors)")
print(f"window: last {WINDOW_H}h -> {len(in_win)} events")
if in_win:
print(f" range: {in_win[0]['ts']} .. {in_win[-1]['ts']}")
print("=" * 78)
# ---- 1. event-type counts ------------------------------------------------
types = Counter(e.get("type", "?") for e in in_win)
print("\n[1] EVENT-TYPE COUNTS")
for t, n in types.most_common():
print(f" {n:6d} {t}")
# ---- per-send assembly ---------------------------------------------------
sends = {}
order = []
def S(e):
sid = e.get("id")
if not sid:
return None
if sid not in sends:
sends[sid] = {"id": sid, "events": [], "first_ts": e["_dt"]}
order.append(sid)
s = sends[sid]
s["events"].append(e)
for k in ("agent", "to", "target"):
if k not in s and e.get(k) is not None:
s[k] = e.get(k)
t = e.get("type")
if t == "verified":
s["verified_seen"] = True
if e.get("thread_uuid"):
s["uuids"] = s.get("uuids", set()) | {e["thread_uuid"]}
if t == "sent":
s["sent_seen"] = True
s["sent_verified"] = bool(e.get("verified"))
if e.get("thread_uuid"):
s["uuids"] = s.get("uuids", set()) | {e["thread_uuid"]}
tags = e.get("tags") or {}
if isinstance(tags, dict) and tags.get("thread"):
s["uuids"] = s.get("uuids", set()) | {tags["thread"]}
if t == "send_done":
s["done"] = True
if t == "sidechat_uuid_capture_failed":
s["capture_failed"] = True
s["capture_out"] = (e.get("out") or "")[:80]
if t == "sidechat_autoprovisioned" and e.get("thread_uuid"):
s["uuids"] = s.get("uuids", set()) | {e["thread_uuid"]}
if t == "alias_resolved" and e.get("thread_uuid"):
s["uuids"] = s.get("uuids", set()) | {e["thread_uuid"]}
if t == "placement_mismatch":
s["placement_mismatch"] = True
if t == "pre_send_assert_failed":
s["gate_failed"] = True
s["gate_reason"] = e.get("reason")
if t == "nav_failed":
s["nav_failed"] = True
s["nav_reason"] = e.get("reason") or "transport"
return s
for e in in_win:
S(e)
started = [sends[i] for i in order
if any(e.get("type") == "send_start" for e in sends[i]["events"])]
print(f"\n[1b] SEND OUTCOMES: {len(started)} sends attempted")
oc = Counter()
for s in started:
if s.get("gate_failed"):
oc["loud-fail: pre_send_assert_failed"] += 1
elif s.get("nav_failed"):
oc["loud-fail: nav_failed"] += 1
elif s.get("sent_seen") and s.get("sent_verified"):
oc["sent verified:true"] += 1
elif s.get("sent_seen"):
oc["sent verified:false"] += 1
elif s.get("done"):
oc["send_done, no sent event"] += 1
else:
oc["abandoned (no terminal event)"] += 1
for k, n in oc.most_common():
print(f" {n:5d} {k}")
print(" (note: main_chat_blocked events are policy blocks, not failures;")
print(" followup_register_failed=signing issues, tangential)")
# ---- 2. sidechat taxonomy ------------------------------------------------
def is_sc(s):
return (s.get("target") or "") not in ("main", "", None)
sc = [s for s in started if is_sc(s)]
print(f"\n[2] SIDECHAT TAXONOMY ({len(sc)} sidechat-targeted sends)")
tax = Counter()
per_target = defaultdict(Counter)
per_pair = defaultdict(Counter)
for s in sc:
if s.get("gate_failed"):
cls = "gate_failed:" + str(s.get("gate_reason"))
elif s.get("nav_failed"):
cls = "nav_failed:" + str(s.get("nav_reason"))
elif s.get("capture_failed"):
cls = "uuid_capture_failed"
elif s.get("placement_mismatch"):
cls = "placement_mismatch"
elif s.get("sent_seen") and s.get("sent_verified"):
cls = ("verified:true, NO uuid anywhere (ghost)"
if not s.get("uuids") else "clean verified")
elif s.get("sent_seen"):
cls = "sent verified:false"
elif s.get("done"):
cls = "done, no sent event"
else:
cls = "abandoned"
tax[cls] += 1
per_target[s.get("target") or "?"][cls] += 1
per_pair[f"{s.get('agent') or '?'}->{s.get('to') or '?'}"][cls] += 1
for k, n in tax.most_common():
print(f" {n:5d} {k}")
print("\n per-target (sends, non-clean, top classes):")
for tgt, c in sorted(per_target.items(), key=lambda x: -sum(x[1].values())):
tot = sum(c.values())
bad = tot - c.get("clean verified", 0)
print(f" {tgt}: {tot} sends, {bad} non-clean {dict(c.most_common(4))}")
print("\n per agent-pair:")
for pair, c in sorted(per_pair.items(), key=lambda x: -sum(x[1].values())):
tot = sum(c.values())
bad = tot - c.get("clean verified", 0)
print(f" {pair}: {tot} sends, {bad} non-clean {dict(c.most_common(4))}")
# ---- 4. ghosts ------------------------------------------------------------
ghosts = [s for s in sc if s.get("sent_seen") and s.get("sent_verified")
and not s.get("uuids")]
print(f"\n[4] GHOST RATE: {len(ghosts)}/{len(sc)} "
f"({100.0 * len(ghosts) / len(sc) if sc else 0:.0f}%) verified:true "
f"sidechat sends with no UUID in any event")
gh = Counter(s["first_ts"].strftime("%m-%d %H") for s in ghosts if s["first_ts"])
print(" ghost hours:", dict(sorted(gh.items())))
print(" proxies: uuid_capture_failed="
f"{sum(1 for s in sc if s.get('capture_failed'))}, "
f"placement_mismatch={sum(1 for s in sc if s.get('placement_mismatch'))}, "
f"pre_send_assert_failed={sum(1 for s in sc if s.get('gate_failed'))}")
# ---- 3. timeline ------------------------------------------------------------
print("\n[3] TIMELINE (hourly, sidechat sends; #=bad, .=ok)")
buckets = defaultdict(Counter)
for s in sc:
if not s.get("first_ts"):
continue
hr = s["first_ts"].strftime("%m-%d %H:00")
buckets[hr]["total"] += 1
bad = not (s.get("sent_seen") and s.get("sent_verified")
and not s.get("capture_failed")
and not s.get("placement_mismatch")
and not s.get("gate_failed") and not s.get("nav_failed"))
if bad:
buckets[hr]["bad"] += 1
for hr in sorted(buckets):
t, b = buckets[hr]["total"], buckets[hr]["bad"]
print(f" {hr} total={t:3d} bad={b:3d} {'#' * b}{'.' * (t - b)}")
wins = defaultdict(Counter)
for s in sc:
if not s.get("first_ts"):
continue
w = s["first_ts"].strftime("%m-%d %H:%M")[:-1] + "0"
wins[w]["total"] += 1
if not (s.get("sent_seen") and s.get("sent_verified")):
wins[w]["bad"] += 1
print(" worst 10-min windows (>=3 sends):")
shown = 0
for w, c in sorted(wins.items(), key=lambda x: -x[1]["bad"]):
if c["total"] >= 3 and shown < 8:
print(f" {w} total={c['total']} bad={c['bad']}")
shown += 1
print(" bad rate by sending node:")
by_node = defaultdict(Counter)
for s in sc:
by_node[s.get("agent") or "?"]["total"] += 1
if not (s.get("sent_seen") and s.get("sent_verified")):
by_node[s.get("agent") or "?"]["bad"] += 1
for node, c in sorted(by_node.items(), key=lambda x: -x[1]["total"]):
r = 100.0 * c["bad"] / c["total"] if c["total"] else 0
print(f" {node}: {c['bad']}/{c['total']} bad ({r:.0f}%)")
# ---- 5. hypothesis evidence ---------------------------------------------------
print("\n[5] HYPOTHESIS EVIDENCE")
nfr = Counter(s.get("nav_reason") for s in sc if s.get("nav_failed"))
print(f" nav_failed reasons: {dict(nfr)}")
land = [s for s in sc if s.get("capture_failed")]
print(f" uuid_capture_failed: {len(land)}, "
f"landing-page outs: {sum(1 for s in land if 'muse.ai/' in (s.get('capture_out') or ''))}")
pairs = Counter()
for e in in_win:
if e.get("type") == "placement_mismatch":
pairs[((e.get("expected_uuid") or "?")[:8],
(e.get("actual_uuid") or "?")[:8], e.get("target"))] += 1
print(" placement_mismatch expected->actual:")
for (a, b, t), n in pairs.most_common(6):
print(f" {n:3d} exp={a}.. act={b}.. target={t}")
ar = Counter(e.get("source") for e in in_win if e.get("type") == "alias_resolved")
print(f" alias_resolved sources: {dict(ar)} (None = field absent, older events)")
ap_try = sum(1 for e in in_win if e.get("type") == "sidechat_autoprovision_start")
ap_ok = sum(1 for e in in_win if e.get("type") == "sidechat_autoprovisioned"
and e.get("thread_uuid"))
print(f" autoprovision: {ap_try} attempts -> {ap_ok} with uuid ({100.0 * ap_ok / ap_try if ap_try else 0:.0f}%)")
rt = Counter()
for s in started:
n = sum(1 for e in s["events"] if e.get("type") == "retry")
if n:
rt[n] += 1
print(f" retry distribution (sends with >=1 retry): {dict(sorted(rt.items()))}")
print("\n DONE.")
if __name__ == "__main__":
sys.exit(main())
-74
View File
@@ -1,74 +0,0 @@
#!/usr/bin/env bash
# dm-sign.sh — produce a signed DM for the fleet DM system.
#
# Signs a message with an SSH key (namespace "dm", Ed25519) and prints the
# signed wire format that `dm.py verify-sig` checks:
#
# [from:<identity>] [id:<id>]
#
# <message body>
#
# -----BEGIN SSH SIGNATURE-----
# ...
# -----END SSH SIGNATURE-----
#
# The signed payload is exactly "[from:X] [id:Y]\n\n<body>" (no trailing
# newline) — verify-sig reconstructs it by stripping everything from the
# signature block onward, so the two must match byte-for-byte.
#
# Usage:
# dm-sign.sh --from <identity> [--key <privkey>] [--id <id>] <message>
#
# Defaults: --key ~/.ssh/id_frontdoor, --id = 8 random hex chars.
# The private key is only ever read locally; it is never moved or copied.
# Pipe the output straight into `dm.py send --raw` (never truncate it).
#
# Example:
# dm.py send --agent opm --target main --raw \
# "$(dm-sign.sh --from operator-main 'hello from the operator')"
set -euo pipefail
FROM=""
KEY="$HOME/.ssh/id_frontdoor"
ID="$(head -c4 /dev/urandom | od -An -tx1 | tr -d ' \n')"
usage() {
sed -n '2,/^set -euo/p' "$0" | sed 's/^# \?//'
}
while [[ $# -gt 0 ]]; do
case "$1" in
--from) FROM="${2:?--from needs a value}"; shift 2 ;;
--key) KEY="${2:?--key needs a value}"; shift 2 ;;
--id) ID="${2:?--id needs a value}"; shift 2 ;;
-h|--help) usage; exit 0 ;;
--) shift; break ;;
-*) echo "error: unknown option: $1" >&2; exit 1 ;;
*) break ;;
esac
done
if [[ $# -eq 0 ]]; then
echo "error: no message given" >&2
echo "usage: dm-sign.sh --from <identity> [--key <privkey>] [--id <id>] <message>" >&2
exit 1
fi
MESSAGE="$*"
[[ -n "$FROM" ]] || { echo "error: --from <identity> is required" >&2; exit 1; }
[[ -f "$KEY" ]] || { echo "error: private key not found: $KEY" >&2; exit 1; }
TD="$(mktemp -d)"
trap 'rm -rf "$TD"' EXIT
PAYLOAD="$TD/payload"
# Payload: header, blank line, body — NO trailing newline (verify-sig strips).
printf '[from:%s] [id:%s]\n\n%s' "$FROM" "$ID" "$MESSAGE" > "$PAYLOAD"
# Never reuse a stale signature: a leftover .sig from an earlier run would
# silently sign the wrong payload (burned 20 minutes on the board, 2026-10-03).
rm -f "$PAYLOAD.sig"
ssh-keygen -Y sign -f "$KEY" -n dm "$PAYLOAD" >/dev/null
printf '[from:%s] [id:%s]\n\n%s\n\n' "$FROM" "$ID" "$MESSAGE"
cat "$PAYLOAD.sig"
-1332
View File
File diff suppressed because it is too large Load Diff
-1
View File
@@ -1 +0,0 @@
/home/super/Projects/NetVM/bin/docs-lookup.py
-606
View File
@@ -1,606 +0,0 @@
#!/usr/bin/env python3
"""docs-lookup.py — Unified Lookup & Regex Passing Tool for docs_internal/.
Provides high-speed queries, regex passing, grammar validation, and surface
lookups for autonomous agents and operators interfacing with NetVM, Box,
and box.muse-dev.online.
Usage:
docs-lookup.py search <query>
docs-lookup.py surfaces [name]
docs-lookup.py sentence [name]
docs-lookup.py regex [name] [--test "<string>"]
docs-lookup.py parse "<string>"
docs-lookup.py get <collection> [key]
docs-lookup.py cli [domain]
docs-lookup.py overview
"""
import argparse
import json
import os
import re
import sys
from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple
NETVM_ROOT = Path(__file__).resolve().parent.parent
LOOKUP_INTERNAL = NETVM_ROOT / "lookup_internal"
if not LOOKUP_INTERNAL.exists() and (NETVM_ROOT / "docs_internal").exists():
LOOKUP_INTERNAL = NETVM_ROOT / "docs_internal"
DOCS_INTERNAL = LOOKUP_INTERNAL
USE_COLOR = sys.stdout.isatty() and os.environ.get("NO_COLOR") is None
def _c(code: str, text: str) -> str:
return f"\033[{code}m{text}\033[0m" if USE_COLOR else str(text)
def c_bold(s: str) -> str: return _c("1", s)
def c_dim(s: str) -> str: return _c("2", s)
def c_green(s: str) -> str: return _c("32", s)
def c_red(s: str) -> str: return _c("31", s)
def c_yellow(s: str) -> str: return _c("33", s)
def c_blue(s: str) -> str: return _c("34", s)
def c_cyan(s: str) -> str: return _c("36", s)
def c_magenta(s: str) -> str: return _c("35", s)
def load_json_file(filename: str) -> Dict[str, Any]:
"""Safely load a JSON file from docs_internal."""
p = DOCS_INTERNAL / filename
if not p.is_file():
return {}
try:
with open(p, "r", encoding="utf-8") as f:
return json.load(f)
except Exception as e:
print(f"Error reading {p}: {e}", file=sys.stderr)
return {}
def load_manifest() -> Dict[str, Any]:
return load_json_file("manifest.json")
def load_sentence_structures() -> Dict[str, Any]:
return load_json_file("sentence_structure.json")
def load_regex_patterns() -> Dict[str, Any]:
return load_json_file("regex_patterns.json")
def load_assistive_surfaces() -> Dict[str, Any]:
return load_json_file("assistive_surfaces.json")
def load_cli_tools() -> Dict[str, Any]:
return load_json_file("cli_tools.json")
def load_fleet_nodes() -> Dict[str, Any]:
return load_json_file("fleet_nodes.json")
def compile_pattern(pat_entry: Dict[str, Any]) -> Tuple[Optional[re.Pattern], Optional[str]]:
"""Compile a regex entry with its configured flags."""
raw_pat = pat_entry.get("pattern", "")
flag_names = pat_entry.get("flags", [])
flags = 0
for fn in flag_names:
if hasattr(re, fn):
flags |= getattr(re, fn)
try:
return re.compile(raw_pat, flags), None
except Exception as e:
return None, str(e)
# ---------------------------------------------------------------------------
# Core Lookup Actions
# ---------------------------------------------------------------------------
def handle_overview(json_mode: bool = False):
manifest = load_manifest()
if json_mode:
print(json.dumps(manifest, indent=2))
return
print(c_bold("\n=== docs_internal — Internal Agent & Operator Lookup Database ==="))
print(c_dim(f"Location: {DOCS_INTERNAL} | Host: {manifest.get('host', 'box.muse-dev.online')}"))
print(f"{manifest.get('description', '')}\n")
print(c_cyan("Available Collections:"))
collections = manifest.get("collections", [])
for col in collections:
cid = col.get("id")
name = col.get("name")
desc = col.get("description")
jfile = col.get("json_file")
mfile = col.get("md_file")
print(f" • {c_bold(cid):<20} {c_green(name)}")
print(f" {c_dim(desc)}")
print(f" {c_dim('Files:')} {c_yellow(jfile)} | {c_yellow(mfile)}")
print()
print(c_dim("Query commands: super docs [search|surfaces|sentence|regex|parse|get|cli]"))
def handle_search(query: str, json_mode: bool = False):
q = query.lower()
results = []
# Search in all JSON files
for jpath in sorted(DOCS_INTERNAL.glob("*.json")):
try:
data = json.loads(jpath.read_text(encoding="utf-8"))
except Exception:
continue
def recurse_search(obj, path=""):
if isinstance(obj, dict):
for k, v in obj.items():
subpath = f"{path}.{k}" if path else k
if q in str(k).lower():
results.append({
"file": jpath.name,
"type": "json_key",
"path": subpath,
"match": str(k),
"preview": str(v)[:160]
})
recurse_search(v, subpath)
elif isinstance(obj, list):
for idx, item in enumerate(obj):
recurse_search(item, f"{path}[{idx}]")
elif isinstance(obj, str):
if q in obj.lower():
results.append({
"file": jpath.name,
"type": "json_value",
"path": path,
"match": obj[:120],
"preview": obj[:240]
})
recurse_search(data)
# Search in Markdown files
for mpath in sorted(DOCS_INTERNAL.glob("*.md")):
try:
content = mpath.read_text(encoding="utf-8")
except Exception:
continue
lines = content.splitlines()
for idx, line in enumerate(lines, 1):
if q in line.lower():
results.append({
"file": mpath.name,
"type": "markdown",
"line": idx,
"match": line.strip(),
"preview": line.strip()
})
if json_mode:
print(json.dumps({"query": query, "count": len(results), "results": results}, indent=2))
return
print(c_bold(f"\nSearch results for '{query}' ({len(results)} matches):"))
if not results:
print(c_dim(" (no matching entries found)"))
return
for r in results[:40]:
if r["type"] == "markdown":
print(f" [{c_yellow(r['file'])}:{c_cyan(str(r['line']))}] {r['match']}")
else:
print(f" [{c_blue(r['file'])}:{c_magenta(r['path'])}] {r['preview']}")
def handle_surfaces(view_name: Optional[str] = None, json_mode: bool = False):
data = load_assistive_surfaces()
views = data.get("views", {})
if view_name:
key = view_name.lower().strip()
v = views.get(key)
if not v:
for k, val in views.items():
if key in k or key in val.get("name", "").lower():
v = val
key = k
break
if not v:
err = {"error": f"Surface '{view_name}' not found", "available": list(views.keys())}
if json_mode:
print(json.dumps(err, indent=2))
else:
print(c_red(f"Error: Surface '{view_name}' not found. Available: {', '.join(views.keys())}"))
sys.exit(1)
if json_mode:
print(json.dumps({key: v}, indent=2))
return
print(c_bold(f"\n=== Assistive Surface: {v.get('name')} (`{key}`) ==="))
print(c_dim(f"Host: {data.get('host')} | Tab ID: {v.get('tab_id')}"))
print(f"{c_cyan('DOM Tab Selector:')} {c_yellow(str(v.get('dom_tab_selector')))}")
print(f"{c_cyan('DOM Pane Selector:')} {c_yellow(str(v.get('dom_pane_selector')))}")
elements = v.get("key_elements", {})
if elements:
print(c_bold("\nKey DOM Elements / Selectors:"))
for el_name, sel in elements.items():
print(f" • {c_magenta(el_name):<20} {c_green(sel)}")
endpoints = v.get("api_endpoints", [])
if endpoints:
print(c_bold("\nAssociated REST API Endpoints:"))
for ep in endpoints:
print(f" • {c_bold(ep.get('method'))} {c_cyan(ep.get('path'))}")
print(f" {c_dim(ep.get('description'))}")
if ep.get("curl_example"):
print(f" {c_dim('curl:')} {c_yellow(ep.get('curl_example'))}")
recipe = v.get("assistive_recipe")
if recipe:
print(c_bold("\nAssistive Recipe:"))
print(f" {recipe}")
print()
return
# All views
if json_mode:
print(json.dumps(data, indent=2))
return
print(c_bold(f"\n=== Assistive Surfaces for {data.get('host', 'box.muse-dev.online')} ==="))
print(c_dim(f"Operator PIN: {data.get('auth', {}).get('pin')} | Session Cookie: {data.get('auth', {}).get('cookie_name')}"))
print()
for k, v in views.items():
name = v.get("name", k)
tab_sel = v.get("dom_tab_selector") or "(modal/overlay)"
eps = [f"{ep.get('method')} {ep.get('path')}" for ep in v.get("api_endpoints", [])]
ep_summary = ", ".join(eps) if eps else "(no direct endpoint)"
print(f" • {c_bold(k):<16} {c_cyan(name):<30} {c_yellow(tab_sel)}")
print(f" {c_dim('API:')} {ep_summary}")
if v.get("assistive_recipe"):
print(f" {c_dim('Hint:')} {v.get('assistive_recipe')[:100]}...")
print()
def handle_sentence(name: Optional[str] = None, json_mode: bool = False):
data = load_sentence_structures()
structs = data.get("structures", {})
if name:
key = name.lower().strip()
s = structs.get(key)
if not s:
for k, val in structs.items():
if key in k or key in val.get("name", "").lower() or key in val.get("protocol_tag", "").lower():
s = val
key = k
break
if not s:
err = {"error": f"Sentence structure '{name}' not found", "available": list(structs.keys())}
if json_mode:
print(json.dumps(err, indent=2))
else:
print(c_red(f"Error: Sentence structure '{name}' not found. Available: {', '.join(structs.keys())}"))
sys.exit(1)
if json_mode:
print(json.dumps({key: s}, indent=2))
return
print(c_bold(f"\n=== Protocol Structure: {s.get('name')} (`{key}`) ==="))
print(f"{c_cyan('Tag:')} {c_bold(s.get('protocol_tag'))}")
print(f"{c_cyan('Template:')} {c_green(s.get('template'))}")
print(f"{c_cyan('Lifecycle:')} {c_yellow(s.get('lifecycle_transition', ''))}")
print(f"\n{c_bold('Description:')}\n {s.get('description')}")
print(f"\n{c_bold('Required Fields:')} {', '.join(s.get('required_fields', []))}")
if s.get("optional_fields"):
print(f"{c_bold('Optional Fields:')} {', '.join(s.get('optional_fields', []))}")
print(f"\n{c_bold('Example:')}\n {c_cyan(s.get('example'))}")
print(f"\n{c_bold('Reply Expectation:')}\n {s.get('reply_expectation')}")
print()
return
if json_mode:
print(json.dumps(data, indent=2))
return
print(c_bold("\n=== Agent Sentence Structures & Conversational Contracts ==="))
print(c_dim("Format rules enforced by response-harvester and self_main_loop.\n"))
for k, s in structs.items():
tag = s.get("protocol_tag", "")
desc = s.get("description", "")
print(f" • {c_bold(k):<18} {c_green(tag):<25} {s.get('name')}")
print(f" {c_dim(desc)}")
print(f" {c_dim('Template:')} {c_cyan(s.get('template', ''))}")
print()
def handle_regex(name: Optional[str] = None, test_str: Optional[str] = None, json_mode: bool = False):
data = load_regex_patterns()
pats = data.get("patterns", {})
if name:
key = name.lower().strip()
p = pats.get(key)
if not p:
for k, val in pats.items():
if key in k or key in val.get("name", "").lower():
p = val
key = k
break
if not p:
err = {"error": f"Regex pattern '{name}' not found", "available": list(pats.keys())}
if json_mode:
print(json.dumps(err, indent=2))
else:
print(c_red(f"Error: Pattern '{name}' not found. Available: {', '.join(pats.keys())}"))
sys.exit(1)
compiled, comp_err = compile_pattern(p)
test_result = None
if test_str is not None:
if compiled:
m = compiled.search(test_str)
test_result = {
"matched": bool(m),
"match_span": m.span() if m else None,
"matched_text": m.group(0) if m else None,
"named_groups": m.groupdict() if m else {},
"groups": list(m.groups()) if m else []
}
else:
test_result = {"matched": False, "error": comp_err}
if json_mode:
out = {key: p, "compiled_ok": bool(compiled)}
if test_str is not None:
out["test_evaluation"] = test_result
print(json.dumps(out, indent=2))
return
print(c_bold(f"\n=== Regex Pattern: {p.get('name')} (`{key}`) ==="))
print(f"{c_cyan('Pattern:')} {c_yellow(p.get('pattern'))}")
print(f"{c_cyan('Flags:')} {', '.join(p.get('flags', [])) or '(none)'}")
print(f"{c_cyan('Description:')} {p.get('description')}")
print(f"{c_cyan('Usage:')} {c_dim(p.get('usage', ''))}")
ng = p.get("named_groups", {})
if ng:
print(c_bold("\nNamed Groups:"))
for gname, gdesc in ng.items():
print(f" • {c_magenta(gname):<16} {gdesc}")
if test_str is not None:
print(c_bold("\nTest Execution Result:"))
print(f" Input: {c_dim(test_str)}")
if test_result.get("matched"):
print(f" Verdict: {c_green('✔ MATCHED')}")
print(f" Matched Text: {c_cyan(test_result['matched_text'])}")
if test_result["named_groups"]:
print(f" Extracted Tokens:")
for k_grp, v_grp in test_result["named_groups"].items():
print(f" - {c_magenta(k_grp)}: {c_yellow(str(v_grp))}")
else:
print(f" Verdict: {c_red('✖ NO MATCH')}")
if comp_err:
print(f" Compile Error: {comp_err}")
print()
return
# List patterns
if json_mode:
print(json.dumps(data, indent=2))
return
print(c_bold("\n=== Master Regex Patterns & Passing Dictionary ==="))
print(c_dim("Patterns tested and calibrated across response-harvester, dm, and loop engines.\n"))
for k, p in pats.items():
print(f" • {c_bold(k):<18} {c_green(p.get('name'))}")
print(f" {c_dim('Pattern:')} {c_yellow(p.get('pattern'))}")
print(f" {c_dim(p.get('description'))}")
print()
def handle_parse(candidate_str: str, json_mode: bool = False):
"""Pass candidate_str through all registered regexes and extract tokens."""
data = load_regex_patterns()
pats = data.get("patterns", {})
matches = []
for k, p in pats.items():
compiled, err = compile_pattern(p)
if not compiled:
continue
m = compiled.search(candidate_str)
if m:
matches.append({
"pattern_key": k,
"pattern_name": p.get("name"),
"matched_text": m.group(0),
"named_groups": m.groupdict(),
"span": m.span()
})
if json_mode:
print(json.dumps({
"input": candidate_str,
"matched_patterns_count": len(matches),
"matches": matches
}, indent=2))
return
print(c_bold(f"\n=== Regex Parse Analysis ==="))
print(f"Input: {c_dim(candidate_str)}\n")
if not matches:
print(c_yellow(" ⚠ No registered regex pattern matched this string."))
return
print(c_green(f"Matched {len(matches)} pattern(s):"))
for match in matches:
print(f"\n • Pattern: {c_bold(match['pattern_name'])} (`{c_cyan(match['pattern_key'])}`)")
print(f" Matched Chunk: {c_yellow(match['matched_text'])}")
ng = match["named_groups"]
if ng:
print(f" Extracted Tokens:")
for gk, gv in ng.items():
print(f" - {c_magenta(gk)}: {c_green(str(gv))}")
print()
def handle_get(collection: str, key: Optional[str] = None, json_mode: bool = False):
col_map = {
"manifest": "manifest.json",
"sentence": "sentence_structure.json",
"sentence_structure": "sentence_structure.json",
"regex": "regex_patterns.json",
"regex_patterns": "regex_patterns.json",
"surfaces": "assistive_surfaces.json",
"assistive_surfaces": "assistive_surfaces.json",
"cli": "cli_tools.json",
"cli_tools": "cli_tools.json",
"fleet": "fleet_nodes.json",
"fleet_nodes": "fleet_nodes.json",
}
fname = col_map.get(collection.lower().strip())
if not fname:
print(c_red(f"Error: Unknown collection '{collection}'. Available: {', '.join(col_map.keys())}"), file=sys.stderr)
sys.exit(1)
data = load_json_file(fname)
if key:
# Check top level or primary container
val = None
for primary in ["structures", "patterns", "views", "domains", "nodes", "collections"]:
if primary in data and isinstance(data[primary], dict) and key in data[primary]:
val = data[primary][key]
break
if val is None and key in data:
val = data[key]
if val is None:
print(c_red(f"Error: Key '{key}' not found in {fname}"), file=sys.stderr)
sys.exit(1)
data = {key: val}
print(json.dumps(data, indent=2))
def handle_cli(domain: Optional[str] = None, json_mode: bool = False):
data = load_cli_tools()
domains = data.get("domains", {})
if domain:
d = domains.get(domain.lower().strip())
if not d:
print(c_red(f"Error: CLI domain '{domain}' not found. Available: {', '.join(domains.keys())}"), file=sys.stderr)
sys.exit(1)
if json_mode:
print(json.dumps({domain: d}, indent=2))
return
print(c_bold(f"\n=== CLI Tool Domain: super {domain} / box {domain} ==="))
print(f"Summary: {d.get('summary')}\n")
for sub in d.get("subcommands", []):
print(f" • {c_green(sub['cmd'])}")
print(f" {c_dim(sub['desc'])}\n")
return
if json_mode:
print(json.dumps(data, indent=2))
return
print(c_bold("\n=== Unified NetVM & Box CLI Tool Catalog ==="))
print(c_dim("Powered by super-cli.py and bin/muse wrapper.\n"))
for dom_k, dom_v in domains.items():
print(f" • {c_bold(dom_k):<16} {c_cyan(dom_v.get('summary'))}")
for sub in dom_v.get("subcommands", [])[:2]:
print(f" - {c_dim(sub['cmd'])}")
print()
# ---------------------------------------------------------------------------
# CLI Argument Parser
# ---------------------------------------------------------------------------
def build_parser():
common = argparse.ArgumentParser(add_help=False)
common.add_argument("--json", action="store_true", help="Output machine-readable JSON")
parser = argparse.ArgumentParser(
description="docs-lookup — Internal Agent & Operator Documentation Query Engine",
parents=[common],
formatter_class=argparse.RawDescriptionHelpFormatter
)
sub = parser.add_subparsers(dest="action")
p_search = sub.add_parser("search", parents=[common], help="Full-text search across docs_internal")
p_search.add_argument("query", help="Search keyword or phrase")
p_surfaces = sub.add_parser("surfaces", parents=[common], help="Lookup assistive surfaces for box.muse-dev.online")
p_surfaces.add_argument("name", nargs="?", default=None, help="Surface or tab name")
p_sentence = sub.add_parser("sentence", parents=[common], help="Lookup agent sentence structures & protocols")
p_sentence.add_argument("name", nargs="?", default=None, help="Structure name (e.g. work_order, result)")
p_regex = sub.add_parser("regex", parents=[common], help="Lookup and test regex patterns")
p_regex.add_argument("name", nargs="?", default=None, help="Pattern name (e.g. work_order, verb)")
p_regex.add_argument("--test", dest="test_str", default=None, help="Test string to evaluate against pattern")
p_parse = sub.add_parser("parse", parents=[common], help="Parse an agent utterance through all regex patterns")
p_parse.add_argument("string", help="String/utterance to parse")
p_get = sub.add_parser("get", parents=[common], help="Query raw JSON collection and key")
p_get.add_argument("collection", help="Collection name (sentence, regex, surfaces, cli, fleet)")
p_get.add_argument("key", nargs="?", default=None, help="Optional specific key")
p_cli = sub.add_parser("cli", parents=[common], help="Lookup CLI commands and syntax")
p_cli.add_argument("domain", nargs="?", default=None, help="CLI domain (fleet, dm, job, etc.)")
p_overview = sub.add_parser("overview", parents=[common], help="Database overview and collections manifest")
return parser
def main():
parser = build_parser()
args = parser.parse_args()
act = args.action
json_mode = getattr(args, "json", False)
if not act or act == "overview":
handle_overview(json_mode)
elif act == "search":
handle_search(args.query, json_mode)
elif act == "surfaces":
handle_surfaces(args.name, json_mode)
elif act == "sentence":
handle_sentence(args.name, json_mode)
elif act == "regex":
handle_regex(args.name, args.test_str, json_mode)
elif act == "parse":
handle_parse(args.string, json_mode)
elif act == "get":
handle_get(args.collection, args.key, json_mode)
elif act == "cli":
handle_cli(args.domain, json_mode)
else:
parser.print_help()
if __name__ == "__main__":
main()
-95
View File
@@ -1,95 +0,0 @@
#!/usr/bin/env bash
# ensure-node-supervision.sh <node> | --all — feed a node to the watchdogs.
#
# Setup (netvm-node-up.sh, hence netvm-provision-node.sh and the onboarding
# pipeline) calls this so every node gets supervision without manual wiring:
# 1. NODES.md registry row (idempotent) — feeds the registry-driven
# supervisors: cdp-relay-watchdog, agent-health.sh, relay-health-check,
# cdp-latency-check. Port from netvm-names pinning (honors
# CDP_PORT_OVERRIDE, so provision's picked port wins when present).
# 2. chromebox-watchdog-<node>.timer unit + enable --now — the one
# supervisor that needs a per-node systemd unit (the @.service
# template already exists). Needs root for the real unit dir.
#
# Env overrides (tests): NODES_MD, UNIT_DIR. systemctl is skipped when
# UNIT_DIR is not the real system dir.
#
# Runs at the end of netvm-node-up.sh (as root); safe to re-run anytime:
# sudo bin/ensure-node-supervision.sh --all
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
NODES_MD="${NODES_MD:-$SCRIPT_DIR/../NODES.md}"
UNIT_DIR="${UNIT_DIR:-/etc/systemd/system}"
# shellcheck disable=SC1091
. "$SCRIPT_DIR/netvm-names.sh"
usage() { echo "usage: ensure-node-supervision.sh <node> | --all" >&2; exit 1; }
ensure_registry_row() {
local node="$1"
if grep -qE "^\|[[:space:]]*$node[[:space:]]*\|" "$NODES_MD" 2>/dev/null; then
echo "registry: $node already in NODES.md"
return 0
fi
netvm_names "$node" || { echo "registry: unknown node $node" >&2; return 1; }
printf '| %s | %s | unknown | %s | active | %s (auto-registered) |\n' \
"$node" "$NETNS" "$CDP_PORT" "$node" >> "$NODES_MD"
echo "registry: added $node (port $CDP_PORT)"
}
ensure_timer() {
local node="$1" unit
unit="$UNIT_DIR/chromebox-watchdog-$node.timer"
if [ -f "$unit" ]; then
echo "timer: chromebox-watchdog-$node.timer already installed"
else
if [ "$UNIT_DIR" = "/etc/systemd/system" ] && [ "$(id -u)" -ne 0 ]; then
echo "timer: need root to install chromebox-watchdog-$node.timer (run with sudo)" >&2
return 1
fi
cat > "$unit" <<EOF
[Unit]
Description=Run chromebox watchdog for $node every 2 minutes
[Timer]
RandomizedDelaySec=30s
OnBootSec=2min
OnUnitActiveSec=2min
Unit=chromebox-watchdog@$node.service
[Install]
WantedBy=timers.target
EOF
echo "timer: installed chromebox-watchdog-$node.timer"
fi
if [ "$UNIT_DIR" = "/etc/systemd/system" ]; then
systemctl daemon-reload
systemctl enable --now "chromebox-watchdog-$node.timer" >/dev/null 2>&1
echo "timer: enabled chromebox-watchdog-$node.timer"
fi
}
ensure_node() {
local node="$1"
ensure_registry_row "$node"
ensure_timer "$node"
}
case "${1:-}" in
--all)
nodes="$(python3 "$SCRIPT_DIR/netvm-registry.py" 2>/dev/null | cut -d: -f1)"
for conf in /etc/netvm/*.conf; do
[ -f "$conf" ] || continue
nodes="$nodes $(basename "$conf" .conf)"
done
seen=""
# shellcheck disable=SC2086 (intended word splitting)
for node in $nodes; do
case " $seen " in *" $node "*) continue;; esac
seen="$seen $node"
ensure_node "$node" || echo "supervision: $node failed (continuing)" >&2
done
;;
""|-h|--help) usage;;
*) ensure_node "$1";;
esac
File diff suppressed because it is too large Load Diff
-377
View File
@@ -1,377 +0,0 @@
#!/usr/bin/env python3
"""
HTTPS exec server for operator remote command execution.
Runs on bl (stable), accepts authenticated POST /exec, returns command output.
Usage:
python3 exec-server.py --port 8443 --token-file /home/super/.exec-token
646 (or any operator) can then:
curl -k -X POST https://100.123.153.75:8443/exec \
-H "Content-Type: application/json" \
-d '{"cmd": "dm.py send --agent 646 --to opm --target main \"hello\"", "token": "..."}'
Via the VM Caddy (no tailnet needed from the agent's container):
curl -k -X POST https://34-139-37-135.sslip.io/exec/exec \
-H "Content-Type: application/json" \
-d '{"cmd": "...", "token": "<per-agent token>"}'
Signature auth (preferred for agents): no secret crosses the wire at all.
The agent signs {"cmd","ts","nonce"} with their registered SSH key
(`ssh-keygen -Y sign -n exec-server`) and posts
{"identity","payload","signature"}. Verified against the signers file
with `ssh-keygen -Y verify`; ts must be within 300s and the nonce unused.
This is the same identity primitive as signed board posts, and it never
trips secret-handling guardrails because a signature is not a secret.
Auth: the master token (TOKEN_FILE) plus per-agent tokens, one file per
agent under TOKEN_DIR (0600). Each agent's token is individually revocable
by deleting its file. The using identity is logged (never the token).
"""
import argparse
import hashlib
import hmac
import json
import secrets
import subprocess
import sys
from http.server import HTTPServer, BaseHTTPRequestHandler
import ssl
# Token file - generated on first run if not exists
TOKEN_FILE = '/home/super/.exec-server-token'
# Per-agent tokens: TOKEN_DIR/<agent> contains that agent's token
TOKEN_DIR = '/home/super/.exec-tokens'
# ssh-keygen signature auth: public keys of fleet identities, one per line
# ("<identity> ssh-ed25519 AAAA..."), synced from the VM's allowed_signers.
SIGNERS_FILE = '/home/super/.exec-signers'
# Seen nonces for replay protection ("<nonce> <ts>" per line).
NONCE_FILE = '/home/super/.exec-nonces'
SIG_NAMESPACE = 'exec-server'
SIG_MAX_SKEW = 300 # seconds; nonces remembered for 2x this
def get_token():
"""Load or generate the auth token."""
try:
with open(TOKEN_FILE, 'r') as f:
return f.read().strip()
except FileNotFoundError:
token = secrets.token_hex(32)
with open(TOKEN_FILE, 'w') as f:
f.write(token)
# Secure permissions
import os
os.chmod(TOKEN_FILE, 0o600)
print(f'Generated new token in {TOKEN_FILE}', file=sys.stderr)
return token
def check_token(token):
"""Check token against the master token and per-agent tokens.
Returns the identity label ('master' or the agent filename), or None."""
if token and hmac.compare_digest(token, get_token()):
return 'master'
import os
try:
names = os.listdir(TOKEN_DIR)
except FileNotFoundError:
return None
for name in names:
p = os.path.join(TOKEN_DIR, name)
if not os.path.isfile(p):
continue
try:
with open(p) as f:
t = f.read().strip()
except OSError:
continue
if t and hmac.compare_digest(token, t):
return name
return None
def check_nonce(nonce):
"""True if the nonce was never used; records it. Prunes expired entries."""
import os
import time
now = time.time()
fresh = []
try:
with open(NONCE_FILE) as f:
for line in f:
parts = line.split()
if len(parts) != 2:
continue
n, t = parts
try:
if now - float(t) < 2 * SIG_MAX_SKEW:
fresh.append((n, t))
except ValueError:
pass
except FileNotFoundError:
pass
if any(n == nonce for n, _ in fresh):
return False
fresh.append((nonce, str(now)))
try:
with open(NONCE_FILE, 'w') as f:
for n, t in fresh:
f.write(f"{n} {t}\n")
os.chmod(NONCE_FILE, 0o600)
except OSError:
return False
return True
def check_signature(identity, payload, signature):
"""Verify an `ssh-keygen -Y` signature over the payload envelope.
Returns the identity on success, None on failure. The envelope must be
JSON {"cmd","ts","nonce"} with a fresh ts and an unused nonce."""
import json as _json
import os
import re
import subprocess
import tempfile
import time
if not re.fullmatch(r'[a-z0-9-]+', identity or ''):
return None
try:
data = _json.loads(payload)
except Exception:
return None
if not isinstance(data, dict):
return None
cmd = data.get('cmd')
ts = data.get('ts')
nonce = data.get('nonce')
if not isinstance(cmd, str) or not cmd:
return None
if not isinstance(nonce, str) or not re.fullmatch(r'[0-9a-fA-F]{16,128}', nonce):
return None
try:
ts = float(ts)
except (TypeError, ValueError):
return None
if abs(time.time() - ts) > SIG_MAX_SKEW:
return None
if not check_nonce(nonce):
return None
sig_path = None
try:
with tempfile.NamedTemporaryFile('w', delete=False, suffix='.sig') as f:
f.write(signature if signature.endswith('\n') else signature + '\n')
sig_path = f.name
p = subprocess.run(
['ssh-keygen', '-Y', 'verify', '-f', SIGNERS_FILE, '-I', identity,
'-n', SIG_NAMESPACE, '-s', sig_path],
input=payload.encode(), capture_output=True, timeout=15)
return identity if p.returncode == 0 else None
except Exception:
return None
finally:
if sig_path:
try:
os.unlink(sig_path)
except OSError:
pass
class ExecHandler(BaseHTTPRequestHandler):
def log_message(self, format, *args):
# Quiet logging, just to stderr
sys.stderr.write(f'{self.client_address[0]} - {format % args}\n')
def do_POST(self):
# Read body
content_length = int(self.headers.get('Content-Length', 0))
if content_length > 1024 * 1024: # 1MB max
self.send_response(413)
self.end_headers()
return
body = self.rfile.read(content_length)
try:
data = json.loads(body)
except json.JSONDecodeError:
self.send_response(400)
self.end_headers()
self.wfile.write(b'{"error": "invalid json"}')
return
if self.path == '/exec/rotate':
self.handle_rotate(data)
return
if self.path != '/exec':
self.send_response(404)
self.end_headers()
return
# Auth: bearer token OR ssh-keygen -Y signature (no secret in transit).
# Bearer: {"token": "...", "cmd": "..."}.
# Signature: {"identity": "...", "payload": "{\"cmd\":...,\"ts\":...,\"nonce\":...}",
# "signature": "<ssh-keygen -Y armor>"}. check_signature
# validates the envelope; cmd comes from the signed payload.
ident = None
cmd = ''
token = data.get('token', '')
if token:
ident = check_token(token)
cmd = data.get('cmd', '')
elif data.get('identity') and data.get('payload') and data.get('signature'):
ident = check_signature(data['identity'], data['payload'], data['signature'])
if ident:
try:
cmd = json.loads(data['payload']).get('cmd', '')
except Exception:
cmd = ''
if not ident:
self.send_response(401)
self.end_headers()
self.wfile.write(b'{"error": "unauthorized"}')
return
sys.stderr.write(f'exec as {ident} from {self.client_address[0]}\n')
# Get command
if not cmd or not isinstance(cmd, str):
self.send_response(400)
self.end_headers()
self.wfile.write(b'{"error": "missing cmd"}')
return
# Execute (with timeout)
timeout = min(data.get('timeout', 60), 300) # max 5 min
try:
result = subprocess.run(
cmd,
shell=True,
capture_output=True,
text=True,
timeout=timeout,
cwd='/home/super'
)
response = {
'stdout': result.stdout,
'stderr': result.stderr,
'rc': result.returncode,
}
except subprocess.TimeoutExpired:
response = {'error': 'timeout', 'rc': -1}
except Exception as e:
response = {'error': str(e), 'rc': -1}
# Send response
resp_body = json.dumps(response).encode()
self.send_response(200)
self.send_header('Content-Type', 'application/json')
self.send_header('Content-Length', str(len(resp_body)))
self.end_headers()
self.wfile.write(resp_body)
def handle_rotate(self, data):
"""POST /exec/rotate - replace an agent's token with a fresh one.
The old token dies immediately; the new one is returned ONLY in the
response body (never logged). Lets an agent bootstrap from a
trust-root-delivered token and end up with one nobody else knows.
Agents may rotate only their own token; master may rotate any
agent's (not its own - that stays manual on bl)."""
import os
import re
token = data.get('token', '')
ident = check_token(token)
if not ident:
self.send_response(401)
self.end_headers()
self.wfile.write(b'{"error": "unauthorized"}')
return
target = data.get('agent', ident)
if not re.fullmatch(r'[a-z0-9-]+', target or ''):
self.send_response(400)
self.end_headers()
self.wfile.write(b'{"error": "bad agent name"}')
return
if target == 'master' or (ident != 'master' and target != ident):
self.send_response(403)
self.end_headers()
self.wfile.write(b'{"error": "forbidden"}')
return
path = os.path.join(TOKEN_DIR, target)
if not os.path.isfile(path):
self.send_response(404)
self.end_headers()
self.wfile.write(b'{"error": "no such agent token"}')
return
new_token = secrets.token_hex(32)
tmp = path + '.tmp'
with open(tmp, 'w') as f:
f.write(new_token + '\n')
os.chmod(tmp, 0o600)
os.replace(tmp, path)
sys.stderr.write(f'token rotated for {target} by {ident}\n')
resp_body = json.dumps({'token': new_token}).encode()
self.send_response(200)
self.send_header('Content-Type', 'application/json')
self.send_header('Content-Length', str(len(resp_body)))
self.end_headers()
self.wfile.write(resp_body)
def do_GET(self):
if self.path == '/health':
self.send_response(200)
self.send_header('Content-Type', 'application/json')
self.end_headers()
self.wfile.write(b'{"status": "ok"}')
else:
self.send_response(404)
self.end_headers()
def main():
global TOKEN_FILE, TOKEN_DIR
parser = argparse.ArgumentParser()
parser.add_argument('--port', type=int, default=8443)
parser.add_argument('--host', default='0.0.0.0')
parser.add_argument('--token-file', default='/home/super/.exec-server-token')
parser.add_argument('--token-dir', default='/home/super/.exec-tokens')
parser.add_argument('--signers-file', default='/home/super/.exec-signers')
parser.add_argument('--nonce-file', default='/home/super/.exec-nonces')
args = parser.parse_args()
TOKEN_FILE = args.token_file
TOKEN_DIR = args.token_dir
global SIGNERS_FILE, NONCE_FILE
SIGNERS_FILE = args.signers_file
NONCE_FILE = args.nonce_file
# Ensure token exists
token = get_token()
print(f'Token: {token[:8]}... (full in {TOKEN_FILE})', file=sys.stderr)
server = HTTPServer((args.host, args.port), ExecHandler)
# Wrap with TLS (self-signed is fine for our use, we use -k)
# Generate self-signed cert if not exists
import os
cert_file = '/home/super/.exec-server-cert.pem'
key_file = '/home/super/.exec-server-key.pem'
if not os.path.exists(cert_file):
print('Generating self-signed cert...', file=sys.stderr)
subprocess.run([
'openssl', 'req', '-x509', '-newkey', 'rsa:2048',
'-keyout', key_file, '-out', cert_file,
'-days', '3650', '-nodes',
'-subj', '/CN=bl-exec-server'
], check=True, capture_output=True)
os.chmod(key_file, 0o600)
context = ssl.SSLContext(ssl.PROTOCOL_TLS_SERVER)
context.load_cert_chain(cert_file, key_file)
server.socket = context.wrap_socket(server.socket, server_side=True)
print(f'Exec server listening on https://{args.host}:{args.port}/exec', file=sys.stderr)
print('Health check: https://<host>:<port>/health', file=sys.stderr)
try:
server.serve_forever()
except KeyboardInterrupt:
print('\nShutting down', file=sys.stderr)
if __name__ == '__main__':
main()
-55
View File
@@ -1,55 +0,0 @@
#!/bin/bash
# exec-sign.sh — call the bl exec-constrained server with SSH-signature auth.
# No bearer token, no secret crosses the wire: you sign the request envelope
# with your registered fleet key and the server verifies it against the
# signers file. A signature is not a secret, so this never trips
# secret-handling guardrails.
#
# Server: exec-constrained.py — named ops ONLY, no arbitrary shell.
# Signature namespace: exec-constrained
# Envelope: {"op","args","ts","nonce"} — op must be in the server allowlist:
# dm.send, dm.thread, dm.read, job.run, chat.messages, chat.send,
# health.check, exec.ping
# Nonce: 16+ hex chars, replay-protected server-side. ts: unix epoch, ±300s skew.
#
# Usage: exec-sign.sh <op> '<args-json>' [identity] [keyfile] [url]
# op a named op, e.g. exec.ping
# args-json JSON object of the op's arguments, e.g. '{}'
# identity defaults to operator-646 (must be a principal in the
# server's signers file)
# keyfile defaults to ~/.ssh/id_frontdoor
# url defaults to https://exec.muse-dev.online/exec (cloudflared).
# NOTE: the old VM Caddy /exec route to bl:8443 died with
# exec-server.py — do not point this at the sslip.io URL.
set -euo pipefail
OP="${1:?usage: exec-sign.sh <op> '<args-json>' [identity] [keyfile] [url]}"
# NOTE: do NOT write this as ${2:-{}} — bash matches the first } as the
# expansion's close brace and appends a literal } when $2 is set.
ARGS_JSON="${2-}"
if [ -z "$ARGS_JSON" ]; then ARGS_JSON='{}'; fi
IDENTITY="${3:-operator-646}"
KEY="${4:-$HOME/.ssh/id_frontdoor}"
URL="${5:-https://exec.muse-dev.online/exec}"
TS=$(date +%s)
NONCE=$(python3 -c "import secrets; print(secrets.token_hex(16))")
PAYLOAD=$(python3 -c "
import json, sys
op, args_json, ts, nonce = sys.argv[1:5]
args = json.loads(args_json)
if not isinstance(args, dict):
raise SystemExit('args-json must be a JSON object')
print(json.dumps({'op': op, 'args': args, 'ts': int(ts), 'nonce': nonce}))
" "$OP" "$ARGS_JSON" "$TS" "$NONCE")
SIG=$(printf '%s' "$PAYLOAD" | ssh-keygen -Y sign -f "$KEY" -n exec-constrained)
BODY=$(python3 -c "
import json, sys
ident, payload, sig = sys.argv[1:4]
print(json.dumps({'identity': ident, 'payload': payload, 'signature': sig}))
" "$IDENTITY" "$PAYLOAD" "$SIG")
# Cloudflare Bot Fight Mode blocks python-urllib POSTs (error 1010):
# use curl with a browser User-Agent instead.
curl -sS -X POST "$URL" \
-H 'Content-Type: application/json' \
-H 'User-Agent: Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36' \
--data "$BODY"
-59
View File
@@ -1,59 +0,0 @@
#!/bin/bash
# exec-watch.sh — 15-min watchdog for the bl exec server's signature-auth path.
# Signs a canary op as the exec-canary identity and POSTs it to the
# local exec-constrained server. Records result in /home/super/.exec-watch-status.json
# and appends to logs/exec-watch.log. Failures are pull-based (status file +
# log); wire push alerting here if the fleet wants paging.
#
# Runs via systemd user timer exec-watch.timer (OnUnitActiveSec=15min).
# Server: exec-constrained.py — named ops ONLY, no arbitrary shell;
# signature namespace: exec-constrained, with {op,args,ts,nonce} envelope.
# NOTE: signers file (/home/super/.exec-signers) is synced MANUALLY from the
# VM's /srv/board/allowed_signers on identity renames/adds — bl cannot ssh
# back to the VM, so there is no pull sync. Operator step, documented in
# docs/TOKEN_POLICY.md.
set -uo pipefail
KEY=/home/super/.exec-canary
URL=https://100.123.153.75:8444/exec
STATUS=/home/super/.exec-watch-status.json
LOG=/home/super/Projects/NetVM/logs/exec-watch.log
ts=$(date +%s)
nonce=$(python3 -c "import secrets; print(secrets.token_hex(16))")
payload=$(python3 -c "import json,sys; print(json.dumps({'op':'exec.ping','args':{},'ts':int(sys.argv[1]),'nonce':sys.argv[2]}))" "$ts" "$nonce")
sig=$(printf '%s' "$payload" | ssh-keygen -Y sign -f "$KEY" -n exec-constrained 2>/dev/null)
if [ -z "${sig:-}" ]; then
result="sign-failed"
else
body=$(python3 -c "import json,sys; print(json.dumps({'identity': 'exec-canary', 'payload': sys.argv[1], 'signature': sys.argv[2]}))" "$payload" "$sig")
out=$(curl -sk -m 25 -X POST "$URL" -H 'Content-Type: application/json' -d "$body" 2>/dev/null)
if echo "$out" | grep -q '"rc": 0'; then
result="ok"
else
result="bad-response"
fi
fi
now_iso=$(date -u +%FT%TZ)
{
python3 - "$STATUS" "$now_iso" "$result" <<'PYEOF'
import json, sys
status_path, now_iso, result = sys.argv[1], sys.argv[2], sys.argv[3]
try:
st = json.load(open(status_path))
except Exception:
st = {}
if result == "ok":
st.update({"last_ok": now_iso, "last_fail": None,
"consecutive_failures": 0, "result": "ok"})
else:
st.update({"last_ok": st.get("last_ok"),
"last_fail": now_iso,
"consecutive_failures": st.get("consecutive_failures", 0) + 1,
"result": result})
json.dump(st, open(status_path, "w"))
print(f"[{now_iso}] exec-watch: {result} "
f"(consecutive_failures={st['consecutive_failures']})")
PYEOF
} >> "$LOG" 2>&1
[ "$result" = "ok" ]
-8
View File
@@ -1,8 +0,0 @@
#!/bin/bash
# flap-check.sh - count browser relaunches per profile in the last hour
# Called by the browser-flap-detector cron. Avoids nested SSH quoting hell.
cutoff=$(date -u -d "1 hour ago" +%Y-%m-%dT%H:%M:%S)
for prof in muse pip 646 opm; do
n=$(grep "\[$prof\]" /home/super/Projects/NetVM/chromebox-watchdog.log 2>/dev/null | grep "relaunch OK" | awk -v d="$cutoff" '{ts=substr($1,2,19); if (ts > d) c++} END {print c+0}')
echo "$prof:$n"
done
-448
View File
@@ -1,448 +0,0 @@
#!/bin/bash
# fleet-alert-check.sh — bl-side critical-condition detector for the fleet alerting pipeline.
#
# Closes the watchdog gap: agent-health.sh logs CRITICAL and restarts browsers,
# but NOTHING pages anyone. This script detects critical conditions, counts
# CONSECUTIVE failures, and emits alert records to an outbox that the
# container-side fleet-alert-relay hook picks up and pages to #lobby.
#
# Checks (bl-side only; VM-side checks live in the container relay):
# cdp:<node> headless Chromium CDP port not listening in the node's netns
# (catches zombie browsers: process alive, CDP not bound)
#
# Paging policy (env-overridable defaults — adjustable, not gates):
# FLEET_ALERT_THRESHOLD=2 consecutive failures before first page (~10 min at 5-min cadence)
# FLEET_ALERT_REALERT_MIN=30 re-page while still critical, at most every 30 min
# FLEET_ALERT_QUIET_HOURS="" e.g. "23:00-07:00" (bl local time); empty = page 24/7.
# First alert for a NEW incident always pages;
# quiet hours only suppress re-pages.
# FLEET_ALERT_DRY_RUN=1 evaluate + print, write no state/outbox, no notify
# FLEET_ALERT_INJECT_FAIL= test hook: comma-separated condition ids to force-fail
# (e.g. FLEET_ALERT_INJECT_FAIL=cdp:pip)
# FLEET_BL_RELAY=1 re-enable the bl-side #lobby relay (default 0/off:
# the container-side hook is the live pager; running
# both double-posts every alert — 2026-10-06)
#
# State: ~/.local/share/fleet-alert/state.json (per-condition consecutive counters)
# Outbox: ~/.local/share/fleet-alert/outbox.jsonl (ALERT/RECOVERY records for the relay)
# Log: /tmp/fleet-alert-check.log
#
# Alert delivery legs:
# 1. outbox record -> container relay -> signed #lobby post (primary page)
# 2. best-effort `box-ctl.py notify` to currently-healthy agents (DM path needs a
# working browser; failures are logged, never fatal)
#
# Installed as user timer fleet-alert-check.timer (every 5 min), mirroring agent-health.timer.
set -uo pipefail
THRESHOLD="${FLEET_ALERT_THRESHOLD:-2}"
REALERT_MIN="${FLEET_ALERT_REALERT_MIN:-30}"
# Approval/input-wait TTLs (seconds): conditions failing longer than this are
# auto-expired (input waits dismissed, key requests denied) instead of paging
# forever. Overridable per environment.
INPUT_WAIT_TTL="${FLEET_ALERT_INPUT_WAIT_TTL:-1800}"
BROWSER_APPROVAL_TTL="${FLEET_ALERT_BROWSER_APPROVAL_TTL:-1800}"
QUIET_HOURS="${FLEET_ALERT_QUIET_HOURS:-}"
DRY_RUN="${FLEET_ALERT_DRY_RUN:-0}"
INJECT_FAIL="${FLEET_ALERT_INJECT_FAIL:-}"
STATE_DIR="${FLEET_ALERT_STATE_DIR:-$HOME/.local/share/fleet-alert}"
BIN="$(cd "$(dirname "$0")" && pwd)"
STATE="$STATE_DIR/state.json"
OUTBOX="$STATE_DIR/outbox.jsonl"
LOG="/tmp/fleet-alert-check.log"
mkdir -p "$STATE_DIR"
[ -f "$STATE" ] || echo '{}' > "$STATE"
touch "$OUTBOX"
NOW=$(date +%s)
log() { echo "$(date -Iseconds) $*" >> "$LOG"; }
# --- shared consecutive-failure state machine (also used by the container relay) ---
# usage: state_machine <cond> <failing 0|1> [ttl_seconds] -> prints "<ACTION> <fails>"
# When ttl_seconds > 0 and the condition has failed longer than the TTL,
# prints "EXPIRED <fails>" so the caller can auto-resolve (dismiss/deny).
# State entries track first_fail_ts (epoch of first consecutive failure).
# ACTION: ALERT_FIRST | ALERT_REALERT | RECOVERY | SUPPRESSED | EXPIRED | NONE
state_machine() {
local cond="$1" failing="$2" ttl="${3:-0}"
THRESHOLD="$THRESHOLD" REALERT_MIN="$REALERT_MIN" QUIET_HOURS="$QUIET_HOURS" \
FLEET_ALERT_DRY_RUN="$DRY_RUN" python3 - "$STATE" "$cond" "$failing" "$ttl" <<'PYEOF'
import json, os, sys, time
state_path, cond, failing_s = sys.argv[1], sys.argv[2], sys.argv[3]
ttl_seconds = int(sys.argv[4]) if len(sys.argv) > 4 else 0
failing = failing_s == "1"
threshold = int(os.environ.get("THRESHOLD", "2"))
realert_min = int(os.environ.get("REALERT_MIN", "30"))
qh = os.environ.get("QUIET_HOURS", "")
dry = os.environ.get("FLEET_ALERT_DRY_RUN") == "1"
now = int(time.time())
def in_quiet(spec):
if not spec:
return False
try:
a, b = spec.split("-")
def m(s):
h, mi = s.split(":")
return int(h) * 60 + int(mi)
cur = time.localtime().tm_hour * 60 + time.localtime().tm_min
s, e = m(a), m(b)
return (s <= cur < e) if s <= e else (cur >= s or cur < e)
except Exception:
return False
try:
st = json.load(open(state_path))
except Exception:
st = {}
e = st.get(cond) or {"fails": 0, "alerted": False, "last_alert_ts": 0}
action = "NONE"
if failing:
if int(e.get("fails", 0)) == 0:
e["first_fail_ts"] = now
e["fails"] = int(e.get("fails", 0)) + 1
# TTL expiry: failing longer than ttl_seconds -> EXPIRED (caller auto-resolves)
if ttl_seconds > 0 and now - int(e.get("first_fail_ts", now)) >= ttl_seconds:
action = "EXPIRED"
# Reset so a fresh incident starts clean after the caller resolves it
e["fails"] = 0
e["alerted"] = False
e.pop("first_fail_ts", None)
else:
due = e["fails"] >= threshold and (
not e.get("alerted") or now - int(e.get("last_alert_ts", 0)) >= realert_min * 60
)
if due:
first = not e.get("alerted")
if not first and in_quiet(qh):
action = "SUPPRESSED"
else:
action = "ALERT_FIRST" if first else "ALERT_REALERT"
e["alerted"] = True
e["last_alert_ts"] = now
else:
if e.get("alerted"):
action = "RECOVERY"
e["fails"] = 0
e["alerted"] = False
e.pop("first_fail_ts", None)
st[cond] = e
if not dry:
json.dump(st, open(state_path, "w"))
print(action, e["fails"])
PYEOF
}
emit_record() { # kind cond detail consecutive
local kind="$1" cond="$2" detail="$3" consec="$4"
local id
id=$(tr -d '-' < /proc/sys/kernel/random/uuid | cut -c1-12)
local rec
rec=$(python3 -c 'import json,sys; print(json.dumps({"id":sys.argv[1],"ts":int(sys.argv[2]),"kind":sys.argv[3],"source":"bl","condition":sys.argv[4],"detail":sys.argv[5],"consecutive":int(sys.argv[6])}))' \
"$id" "$NOW" "$kind" "$cond" "$detail" "$consec")
if [ "$DRY_RUN" = "1" ]; then
log "DRY-RUN would emit: $rec"
else
echo "$rec" >> "$OUTBOX"
log "emitted $kind $cond x$consec"
fi
}
box_notify() {
# Best-effort DM to healthy agents via box-ctl. Never fatal.
# Fan-out runs in parallel with a per-notify timeout so one hung DM
# path can't stall the 5-minute check loop.
local cond="$1"
local detail="$2"
local msg="[fleet-alert] CRITICAL ${cond}: ${detail}"
msg="${msg:0:240}"
local agent pids=""
for agent in $HEALTHY_AGENTS; do
if [ "$DRY_RUN" = "1" ]; then
log "DRY-RUN would notify $agent"
continue
fi
(
if timeout 60 python3 "$BIN/box-ctl.py" notify "$agent" "$msg" >/dev/null 2>&1; then
log "notified $agent re $cond"
else
log "notify $agent failed re $cond (best-effort)"
fi
) &
pids="$pids $!"
done
local p
for p in $pids; do wait "$p" 2>/dev/null; done
}
notify_input_wait() {
# Targeted DM for input waits (2026-10-05): DM ONLY the specific agent
# whose session is waiting for human input -- not a broadcast to all
# healthy agents. The #lobby post still fires via the relay leg for
# human visibility; this DM ensures the responsible operator sees it
# in their sidechat without digging through lobby noise.
# Best-effort: never fatal to the 5-minute check loop.
local node="$1"
local detail="$2"
local msg="[fleet-alert] INPUT WAIT: ${detail} -- reply: box approval reply ${node} \"<msg>\" or box notify ${node} \"<msg>\""
msg="${msg:0:900}"
if [ "$DRY_RUN" = "1" ]; then
log "DRY-RUN would DM $node re input_wait"
return 0
fi
if timeout 60 python3 "$BIN/box-ctl.py" notify "$node" "$msg" >/dev/null 2>&1; then
log "input_wait targeted DM sent to $node"
else
log "input_wait DM to $node failed (best-effort, non-fatal)"
fi
}
injected() { # cond -> 0 if injected-fail
case ",$INJECT_FAIL," in *,"$1,"*) return 0;; *) return 1;; esac
}
HEALTHY_AGENTS=""
"$BIN/netvm-registry.py" 2>/dev/null | while IFS=: read -r node port; do
[ -n "$node" ] && [ -n "$port" ] || continue
cond="cdp:$node"
procs=$(pgrep -f "chromium.*profiles/$node" 2>/dev/null | wc -l)
if sudo -n ip netns exec "warp-$node" ss -tln 2>/dev/null | grep -q ":$port "; then
failing=0
HEALTHY_AGENTS="$HEALTHY_AGENTS $node"
else
failing=1
fi
injected "$cond" && failing=1
detail="CDP $port not listening in netns warp-$node (chromium procs=$procs)"
# NOTE: HEALTHY_AGENTS set inside the pipeline subshell is lost; recompute below.
read -r action fails < <(state_machine "$cond" "$failing")
case "$action" in
ALERT_FIRST|ALERT_REALERT)
emit_record "ALERT" "$cond" "$detail" "$fails"
echo "$cond" >> "$STATE_DIR/.alerts.tmp"
;;
RECOVERY)
emit_record "RECOVERY" "$cond" "CDP $port listening again in netns warp-$node" "$fails"
;;
SUPPRESSED)
log "$cond still critical x$fails — re-page suppressed by quiet hours ($QUIET_HOURS)"
;;
esac
done
# --- Agent approval blockage & input wait detection ---
"$BIN/netvm-registry.py" 2>/dev/null | while IFS=: read -r node port; do
[ -n "$node" ] || continue
node_data=$(python3 -c "
import sys, json
sys.path.insert(0, '$BIN')
import approvals
info = approvals.inspect_node_approvals('$node')
out = {
'has_pending': info.get('has_pending', False),
'target': info.get('target') or info.get('ip') or 'unknown',
'title': info.get('title') or '',
'waits': info.get('input_waits') or []
}
print(json.dumps(out))
" 2>/dev/null || echo '{"has_pending":false,"target":"unknown","title":"","waits":[]}')
# 1. Egress permission dialog
cond="approval:$node"
has_pending=$(python3 -c "import json,sys; print(1 if json.loads(sys.argv[1]).get('has_pending') else 0)" "$node_data" 2>/dev/null || echo 0)
if [ "$has_pending" = "1" ]; then
failing=1
target=$(python3 -c "import json,sys; print(json.loads(sys.argv[1]).get('target','unknown'))" "$node_data" 2>/dev/null || echo unknown)
detail="Agent $node held up on browser approval for $target"
else
failing=0
detail="Agent $node approvals clear"
fi
injected "$cond" && failing=1
read -r action fails < <(state_machine "$cond" "$failing" "$BROWSER_APPROVAL_TTL")
case "$action" in
ALERT_FIRST|ALERT_REALERT)
emit_record "ALERT" "$cond" "$detail" "$fails"
echo "$cond" >> "$STATE_DIR/.alerts.tmp"
;;
RECOVERY)
emit_record "RECOVERY" "$cond" "$detail" "$fails"
;;
SUPPRESSED)
log "$cond still critical x$fails — re-page suppressed"
;;
EXPIRED)
# Browser approval dialog exceeded BROWSER_APPROVAL_TTL without a
# human decision: fail closed by denying it.
log "$cond EXPIRED after ${BROWSER_APPROVAL_TTL}s without human decision — auto-denying (fail closed)"
python3 - "$node" <<'PYEOF3'
import sys, json
sys.path.insert(0, "/home/super/Projects/NetVM/bin")
import approvals
node = sys.argv[1]
print(json.dumps(approvals.deny_node_approval(node, caller="approval-ttl-expire")))
approvals.log_box_ctl("approval-expired", name=node, caller="approval-ttl-expire",
extra={"note": "browser approval TTL elapsed; auto-denied (fail closed)"})
PYEOF3
emit_record "RECOVERY" "$cond" "Agent $node browser approval expired after ${BROWSER_APPROVAL_TTL}s; auto-denied" "$fails"
;;
esac
# 2. Sidebar task waiting on human input
cond_in="input_wait:$node"
wait_summary=$(python3 -c "
import json,sys
w = json.loads(sys.argv[1]).get('waits', [])
if w:
print('; '.join(f\"{item.get('task')}: {item.get('status')}\" for item in w)[:120])
" "$node_data" 2>/dev/null || true)
if [ -n "$wait_summary" ]; then
failing_in=1
detail_in="Agent $node task waiting for human input: $wait_summary"
else
failing_in=0
detail_in="Agent $node tasks running"
fi
injected "$cond_in" && failing_in=1
read -r action_in fails_in < <(state_machine "$cond_in" "$failing_in" "$INPUT_WAIT_TTL")
case "$action_in" in
ALERT_FIRST|ALERT_REALERT)
emit_record "ALERT" "$cond_in" "$detail_in" "$fails_in"
echo "$cond_in|$detail_in" >> "$STATE_DIR/.alerts.tmp"
;;
RECOVERY)
emit_record "RECOVERY" "$cond_in" "$detail_in" "$fails_in"
;;
SUPPRESSED)
log "$cond_in still critical x$fails_in — re-page suppressed"
;;
EXPIRED)
# Input wait exceeded INPUT_WAIT_TTL without human response:
# auto-dismiss so the agent unblocks. Log the expiry and emit a
# RECOVERY record (the wait is gone, not merely un-paged).
log "$cond_in EXPIRED after ${INPUT_WAIT_TTL}s without human input — auto-dismissing"
python3 - "$node" <<'PYEOF2'
import sys
sys.path.insert(0, "/home/super/Projects/NetVM/bin")
import approvals, json
node = sys.argv[1]
info = approvals.inspect_node_approvals(node)
for w in info.get("input_waits", []) or []:
t = w.get("task")
if t:
approvals.mark_wait_responded(node, t, caller="approval-ttl-expire")
approvals.log_box_ctl("approval-wait-expired", name=node, caller="approval-ttl-expire",
extra={"note": "input wait TTL elapsed; auto-dismissed"})
print(json.dumps(approvals.dismiss_node_task(node, caller="approval-ttl-expire")))
PYEOF2
emit_record "RECOVERY" "$cond_in" "Agent $node input wait expired after ${INPUT_WAIT_TTL}s; auto-dismissed" "$fails_in"
;;
esac
done
# --- Warp partition detection (2026-10-04) ---
# A partitioned node has a live browser + CDP but no internet egress: the
# chromebox watchdog sees a healthy browser while all automation fails.
# Condition id: partition:<node>. The detail names the node, the WireGuard
# handshake age, the egress probe result, and the timestamp, and says
# PARTITION explicitly so #lobby readers can tell a network partition from
# a browser crash at a glance. Anti-spam comes from the shared consecutive-
# failure state machine (2 consecutive failures before first page, re-page
# at most every 30 min).
WARP_PROBE_URL="${WARP_PROBE_URL:-https://1.1.1.1/cdn-cgi/trace}"
warp_partition_probe() { # <node> -> prints "<handshake_age_s|unknown> <ok|FAIL>"
local node="$1" iface epoch now age_s
iface=$(sudo -n ip netns exec "warp-$node" sh -c 'wg show interfaces 2>/dev/null | head -1')
now=$(date +%s)
epoch=$(sudo -n ip netns exec "warp-$node" wg show "$iface" latest-handshakes 2>/dev/null | awk '{print $2}')
case "$epoch" in ''|*[!0-9]*) age_s="unknown" ;; *) age_s=$(( now - epoch )) ;; esac
if sudo -n ip netns exec "warp-$node" curl -s -m 8 -o /dev/null "$WARP_PROBE_URL" 2>/dev/null; then
echo "$age_s ok"
else
echo "$age_s FAIL"
fi
}
"$BIN/netvm-registry.py" 2>/dev/null | while IFS=: read -r node port; do
[ -n "$node" ] && [ -n "$port" ] || continue
cond="partition:$node"
read -r hs_age probe_res < <(warp_partition_probe "$node")
ts=$(date -u +%FT%TZ)
if [ "$probe_res" = "ok" ]; then
failing=0
detail="warp egress restored for $node at $ts (probe $WARP_PROBE_URL ok)"
else
failing=1
detail="PARTITION $node: warp egress down at $ts (handshake ${hs_age}s ago, probe $WARP_PROBE_URL FAILED)"
fi
if injected "$cond"; then
failing=1
detail="PARTITION $node: warp egress down at $ts (handshake ${hs_age}s ago, probe $WARP_PROBE_URL FAILED) [INJECTED]"
fi
read -r action fails < <(state_machine "$cond" "$failing")
case "$action" in
ALERT_FIRST|ALERT_REALERT)
emit_record "ALERT" "$cond" "$detail" "$fails"
echo "$cond" >> "$STATE_DIR/.alerts.tmp"
;;
RECOVERY)
emit_record "RECOVERY" "$cond" "$detail" "$fails"
;;
SUPPRESSED)
log "$cond still critical x$fails — re-page suppressed by quiet hours ($QUIET_HOURS)"
;;
esac
done
# Recompute healthy agents in the main shell (pipeline subshell above can't export).
HEALTHY_AGENTS=""
"$BIN/netvm-registry.py" 2>/dev/null | while IFS=: read -r node port; do
[ -n "$node" ] && [ -n "$port" ] || continue
if sudo -n ip netns exec "warp-$node" ss -tln 2>/dev/null | grep -q ":$port "; then
echo "$node"
fi
done > "$STATE_DIR/.healthy.tmp"
HEALTHY_AGENTS=$(tr '\n' ' ' < "$STATE_DIR/.healthy.tmp")
rm -f "$STATE_DIR/.healthy.tmp"
# Notify for this run's alerts (best effort). Skip entirely when nothing is healthy
# (notify needs a working browser via dm.py) or in dry-run.
if [ -n "$HEALTHY_AGENTS" ] && [ -f "$STATE_DIR/.alerts.tmp" ]; then
while IFS= read -r line; do
# alerts.tmp format: "cond" or "cond|detail" (input_wait carries detail)
cond="${line%%|*}"
detail="${line#*|}"
[ "$detail" = "$line" ] && detail=""
[ -n "$cond" ] || continue
case "$cond" in
input_wait:*)
# Targeted: DM only the waiting agent, not a broadcast.
node="${cond#input_wait:}"
if [ -n "$detail" ]; then
notify_input_wait "$node" "$detail"
else
box_notify "$cond" "see #lobby for detail"
fi
;;
*)
box_notify "$cond" "see #lobby for detail"
;;
esac
done < "$STATE_DIR/.alerts.tmp"
elif [ -f "$STATE_DIR/.alerts.tmp" ]; then
log "no healthy agents — box notify skipped (DM path needs a working browser)"
fi
rm -f "$STATE_DIR/.alerts.tmp"
tail -500 "$LOG" > "$LOG.tmp" 2>/dev/null && mv "$LOG.tmp" "$LOG"
log "check complete"
# Bl-side #lobby relay: DISABLED by default (FLEET_BL_RELAY=1 to re-enable).
# The container-side hook is the live pager; the bl relay never successfully
# posted (missing CHAT_KEYFILE) and enabling it now would double-post every
# alert in a second format. Re-enable only alongside retiring the container
# hook (and per the relay header, with opm sign-off).
if [ "${FLEET_BL_RELAY:-0}" = "1" ] && [ "$DRY_RUN" -eq 0 ] && [ -x "$BIN/fleet-alert-relay.sh" ]; then
"$BIN/fleet-alert-relay.sh" >> "$LOG" 2>&1 || true
fi
-336
View File
@@ -1,336 +0,0 @@
#!/bin/bash
# bin/fleet-alert-relay.sh — idempotent relay: fleet-alert outbox.jsonl -> #lobby.
#
# PROPOSAL ONLY (board ticket ae28ac8b735f). No bl infra is touched by this
# branch; deployment needs opm review + sign-off. See
# docs/FLEET-ALERT-DUP-POST-GATE.md.
#
# NOTE (2026-10-06): auto-invoke from fleet-alert-check.sh is disabled by
# default (FLEET_BL_RELAY=1 re-enables). The container-side hook pages
# #lobby today; do not re-enable without retiring it first.
#
# The 2026-10-05 11:28Z incident: one RECOVERY record in the outbox became two
# identical verified #lobby posts (seq 642/643, 3.35s apart) because the relay
# leg had no idempotency: append-only outbox, no consume tracking, no content
# dedup, and the chat POST API has no idempotency key.
#
# Gates implemented here:
# 1. posted-watermark — each posted outbox record id is appended to
# posted.log; records already posted are skipped.
# 2. content-hash + 10-min TTL — sha256 of the formatted alert text in a
# file-backed seen-set (seen-hashes.log); identical text re-posts inside
# the window are dropped. Catches same-record re-posts AND identical
# text from different records.
# 3. retry discipline — never blind-retry a chat POST after a timeout /
# phantom-000. Read the #lobby tail first; re-post only if absent.
#
# Usage:
# fleet-alert-relay.sh [--dry-run] [--self-test]
#
# Env: FLEET_ALERT_DIR (default ~/.local/share/fleet-alert),
# CHAT_BASE (default https://chat.muse-dev.online),
# CHAT_IDENTITY (default operator-646), CHAT_KEYFILE (default ~/.ssh/id_frontdoor)
set -euo pipefail
ALERT_DIR="${FLEET_ALERT_DIR:-$HOME/.local/share/fleet-alert}"
OUTBOX="$ALERT_DIR/outbox.jsonl"
POSTED="$ALERT_DIR/posted.log"
SEEN="$ALERT_DIR/seen-hashes.log"
LOCKF="$ALERT_DIR/relay.lock"
CHAT_BASE="${CHAT_BASE:-https://chat.muse-dev.online}"
API="$CHAT_BASE/api/chat"
IDENTITY="${CHAT_IDENTITY:-operator-646}"
KEYFILE="${CHAT_KEYFILE:-$HOME/.ssh/id_frontdoor}"
CHANNEL="#lobby"
TTL_SECS=600
UA='Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0 Safari/537.36'
DRY_RUN=0
log() { echo "[fleet-alert-relay] $*" >&2; }
# ---------------------------------------------------------------- transport
# Factored so --self-test can stub them. Return contracts:
# transport_tail -> prints "<ts> <identity> <message>" lines for recent #lobby (one per line, tab-safe: message may contain spaces, fields split on first two spaces)
# transport_post "<text>" -> prints "OK <msg-id>" or "UNKNOWN <reason>"; exit 0 always (callers decide)
transport_tail() {
curl -s -m 15 -A "$UA" "$API/history?channel=%23lobby&limit=50" \
| python3 -c '
import json,sys
try:
d=json.load(sys.stdin)
except Exception:
sys.exit(0)
for m in d.get("messages",[]):
msg=(m.get("message") or "").replace(chr(10)," ")
print(m.get("ts"), m.get("identity"), msg)
'
}
transport_post() { # $1 = text
local text="$1" ts sig payload resp http
ts="$(date +%s)"
if ! sig="$(sign_payload "$(printf '%s\n%s\n%s' "$ts" "$CHANNEL" "$text")")"; then
echo "UNKNOWN sign-failed($KEYFILE)"; return 0
fi
payload="$(MSG="$text" TS="$ts" SIG="$sig" python3 -c '
import json,os
print(json.dumps({"channel":os.environ["CHANNEL"],"identity":os.environ["IDENTITY"],
"message":os.environ["MSG"],"ts":int(os.environ["TS"]),"signature":os.environ["SIG"]}))')"
resp="$(curl -s -m 20 -A "$UA" -X POST -H 'Content-Type: application/json' \
-d "$payload" -w '\n%{http_code}' "$API/post" 2>/dev/null || echo -e '\n000')"
http="$(printf '%s' "$resp" | tail -n1)"
resp="$(printf '%s' "$resp" | sed '$d')"
case "$http" in
2*)
local mid
mid="$(printf '%s' "$resp" | python3 -c '
import json,sys
try:
d=json.load(sys.stdin); print(d.get("id","") if d.get("ok") else "")
except Exception:
print("")')"
if [ -n "$mid" ]; then echo "OK $mid"; else echo "UNKNOWN bad-body-$http"; fi
;;
*) echo "UNKNOWN http-$http" ;;
esac
return 0
}
# file-based signing (never pipe): stdin-piped signatures intermittently
# fail verification (2026-10-02 incident).
sign_payload() { # $1 = payload -> armored signature
local tmpd pf
tmpd="$(mktemp -d)" || return 1
pf="$tmpd/payload"
printf '%s' "$1" > "$pf"
rm -f "$pf.sig"
ssh-keygen -Y sign -f "$KEYFILE" -n chat "$pf" </dev/null >/dev/null 2>&1 \
|| { rm -rf "$tmpd"; return 1; }
cat "$pf.sig"
rm -rf "$tmpd"
}
# ----------------------------------------------------------------- formatting
# Canonical text for one outbox record (JSON on stdin):
# ALERT -> [fleet-alert] CRITICAL: <node> warp PARTITION x<n>
# RECOVERY -> [fleet-alert] RECOVERED: <node> warp PARTITION
# condition is "<type>:<node>" (partition|cdp).
format_text() {
python3 -c '
import json,sys
r=json.load(sys.stdin)
cond=r.get("condition","")
typ,_,node=cond.partition(":")
label={"partition":"PARTITION","cdp":"CDP"}.get(typ,typ.upper())
kind=r.get("kind","")
if kind=="ALERT":
n=r.get("consecutive",r.get("count","?"))
print("[fleet-alert] CRITICAL: %s warp %s x%s" % (node,label,n))
elif kind=="RECOVERY":
print("[fleet-alert] RECOVERED: %s warp %s" % (node,label))
else:
print("[fleet-alert] %s: %s warp %s" % (kind,node,label))
'
}
content_hash() { printf '%s' "$1" | sha256sum | cut -d' ' -f1; }
# ------------------------------------------------------------- gate 1: watermark
already_posted() { # $1 = record id
[ -f "$POSTED" ] && grep -qxF "$1" "$POSTED"
}
mark_posted() { # $1 = record id
printf '%s\n' "$1" >> "$POSTED"
}
# ------------------------------------------------------- gate 2: content-hash TTL
seen_recently() { # $1 = hash -> 0 if seen within TTL
local h="$1" now ts
[ -f "$SEEN" ] || return 1
now="$(date +%s)"
while read -r h2 ts _rest; do
[ "$h2" = "$h" ] || continue
if [ "$(( now - ts ))" -lt "$TTL_SECS" ]; then return 0; fi
done < "$SEEN"
return 1
}
mark_seen() { # $1 = hash
printf '%s %s\n' "$1" "$(date +%s)" >> "$SEEN"
}
prune_seen() {
local now tmp
[ -f "$SEEN" ] || return 0
now="$(date +%s)"; tmp="$(mktemp)"
while read -r h ts _rest; do
[ -n "$h" ] && [ "$(( now - ts ))" -lt "$TTL_SECS" ] && printf '%s %s\n' "$h" "$ts" >> "$tmp"
done < "$SEEN"
mv "$tmp" "$SEEN"
}
# ------------------------------------------------- gate 3: check-before-(re)post
# 0 if the exact text appears in the recent #lobby tail.
lobby_has_text() { # $1 = text
local want="$1"
# transport_tail prints "<ts> <identity> <message>" per line; the message
# itself may contain spaces, so strip only the first two fields.
transport_tail | python3 -c '
import sys
want=sys.argv[1]
for line in sys.stdin:
parts=line.rstrip("\n").split(" ",2)
if len(parts)==3 and parts[2]==want:
sys.exit(0)
sys.exit(1)' "$want"
}
relay_record() { # $1 = record id, $2 = formatted text
local rid="$1" text="$2" h out
h="$(content_hash "$text")"
# Gate 1: posted watermark
if already_posted "$rid"; then
log "skip $rid: already in posted.log"
return 0
fi
# Gate 2: content hash within TTL
if seen_recently "$h"; then
log "skip $rid: identical text posted <10min ago (hash $h)"
mark_posted "$rid" # consume it so it never retries later
return 0
fi
# Gate 3a: somebody else already posted it (covers agent-driven relay overlap)
if lobby_has_text "$text"; then
log "skip $rid: text already in #lobby tail"
mark_posted "$rid"; mark_seen "$h"
return 0
fi
if [ "$DRY_RUN" = 1 ]; then
log "dry-run: would post $rid: $text"
return 0
fi
# Attempt one POST. Any ambiguous outcome -> verify via tail, never blind-retry.
out="$(transport_post "$text")"
case "$out" in
OK*)
log "posted $rid -> ${out#OK }"
mark_posted "$rid"; mark_seen "$h"
return 0
;;
UNKNOWN*)
log "post $rid outcome UNKNOWN (${out#UNKNOWN }); checking #lobby tail before any retry"
sleep 2
if lobby_has_text "$text"; then
log "post $rid landed despite ${out#UNKNOWN } — marking posted, no retry"
mark_posted "$rid"; mark_seen "$h"
return 0
fi
log "post $rid absent from tail — ONE retry only"
out="$(transport_post "$text")"
case "$out" in
OK*)
log "retry posted $rid -> ${out#OK }"
mark_posted "$rid"; mark_seen "$h"
return 0
;;
*)
log "ERROR: post $rid still ${out%% *} after one retry; leaving unposted for next run (tail-check will catch it if it landed)"
return 1
;;
esac
;;
esac
}
main() {
local rid rec text fails=0
[ -f "$OUTBOX" ] || { log "no outbox ($OUTBOX); nothing to do"; return 0; }
mkdir -p "$ALERT_DIR"
touch "$POSTED" "$SEEN"
prune_seen
while IFS= read -r rec; do
[ -n "$rec" ] || continue
rid="$(printf '%s' "$rec" | python3 -c 'import json,sys; print(json.load(sys.stdin).get("id",""))')"
[ -n "$rid" ] || { log "skip record with no id"; continue; }
text="$(printf '%s' "$rec" | format_text)"
relay_record "$rid" "$text" || fails=$((fails+1))
done < "$OUTBOX"
return "$fails"
}
# ------------------------------------------------------------------ self-test
# Acceptance: two identical alert submissions <10 min apart -> exactly one
# #lobby post. Stubs the transport; exercises all three gates.
# NOTE: transport_post is invoked via command substitution (subshell), so the
# stub counts calls with a file, not a variable.
self_test() {
local td calls lobby ok=1 n sk sig_out old_key
td="$(mktemp -d)"; export FLEET_ALERT_DIR="$td"
ALERT_DIR="$td"; OUTBOX="$td/outbox.jsonl"; POSTED="$td/posted.log"
SEEN="$td/seen-hashes.log"; LOCKF="$td/relay.lock"
calls="$td/calls.log"; lobby="$td/lobby.log"
touch "$calls" "$lobby"
# sign_payload must round-trip with a valid key and fail cleanly without
# one (2026-10-06: missing ~/.ssh/id_frontdoor broke every #lobby post
# with an undiagnosable bare "sign-failed").
old_key="$KEYFILE"
sk="$td/signkey"
ssh-keygen -t ed25519 -f "$sk" -N '' -q >/dev/null 2>&1 \
|| { echo "FAIL: cannot generate ephemeral test key"; ok=0; }
if KEYFILE="$sk" sig_out="$(sign_payload "self-test")"; then
case "$sig_out" in
*"BEGIN SSH SIGNATURE"*) : ;;
*) echo "FAIL: sign_payload output not armored"; ok=0 ;;
esac
else
echo "FAIL: sign_payload failed with a valid key"; ok=0
fi
if KEYFILE="$td/no-such-key" sign_payload "self-test" >/dev/null 2>&1; then
echo "FAIL: sign_payload succeeded with a missing key"; ok=0
fi
KEYFILE="$old_key"
# Two identical submissions: same text, different record ids (the 11:28Z shape)
printf '%s\n' \
'{"id":"rec-A","ts":1791199616,"kind":"RECOVERY","condition":"partition:def"}' \
'{"id":"rec-B","ts":1791199617,"kind":"RECOVERY","condition":"partition:def"}' \
> "$OUTBOX"
transport_post() { # stub: 1st call "times out" but the server DID accept it
echo "call" >> "$calls"
printf '%s\n' "$1" >> "$lobby"
n=$(wc -l < "$calls")
if [ "$n" = 1 ]; then echo "UNKNOWN phantom-000"; else echo "OK stub-$n"; fi
return 0
}
transport_tail() { # stub: the lobby shows whatever "landed"
while IFS= read -r m; do printf '%s opm %s\n' "$(date +%s)" "$m"; done < "$lobby"
}
main
n=$(wc -l < "$calls")
[ "$n" -eq 1 ] || { echo "FAIL: expected exactly 1 post call, got $n"; ok=0; }
grep -qx "rec-A" "$POSTED" || { echo "FAIL: rec-A not watermarked"; ok=0; }
grep -qx "rec-B" "$POSTED" || { echo "FAIL: rec-B not watermarked (gate 2 should consume it)"; ok=0; }
# Identical re-submission (new record id, same text) must not post again
printf '%s\n' '{"id":"rec-C","ts":1791199700,"kind":"RECOVERY","condition":"partition:def"}' >> "$OUTBOX"
main
n=$(wc -l < "$calls")
[ "$n" -eq 1 ] || { echo "FAIL: identical re-submission posted again (calls=$n)"; ok=0; }
rm -rf "$td"
[ "$ok" = 1 ] && echo "SELF-TEST PASS: 2 identical submissions -> 1 post call; re-submission -> 0 new calls" && return 0
return 1
}
case "${1:-}" in
--dry-run) DRY_RUN=1; shift ;;
--self-test) self_test; exit $? ;;
esac
# Single-flight the relay (defense in depth; the detector's own flock is separate).
exec 9>"$LOCKF"
if ! flock -n 9; then
log "another relay run holds the lock; exiting"
exit 0
fi
main
-12
View File
@@ -1,12 +0,0 @@
#!/bin/bash
# fleet-status.sh - one-line health per profile: browser up/down + relay code
declare -A ports=( [muse]=9410 [pip]=9420 [646]=9430 [opm]=9440 )
declare -A veth=( [muse]=10.201.35.2 [pip]=10.201.87.2 [646]=10.201.202.2 [opm]=10.201.157.2 )
for p in muse pip 646 opm; do
port=${ports[$p]}
if pgrep -f "remote-debugging-port=$port" >/dev/null 2>&1; then b=UP; else b=DOWN; fi
r=$(curl -s -m 4 -o /dev/null -w "%{http_code}" "http://${veth[$p]}:$port/json/version" 2>/dev/null || echo 000)
echo "$p: browser=$b relay=$r"
done
echo ===
tail -20 /home/super/Projects/NetVM/chromebox-watchdog.log | grep -E "FAILED|unhealthy" | tail -4
-578
View File
@@ -1,578 +0,0 @@
#!/usr/bin/env python3
"""
followup-sweeper.py — Autonomous follow-up deadline tracking and nudge sweeper.
Monitors pending follow-ups in followups.json, delivers progressive nudges to
recipients when deadlines expire (in-thread first, Main Chat on final nudge),
and executes terminal escalations to opm when all nudges are exhausted.
Usage:
python3 followup-sweeper.py --once
python3 followup-sweeper.py --loop --interval 30
"""
import argparse
import json
import os
import re
import subprocess
import sys
import time
from datetime import datetime, timezone, timedelta
from pathlib import Path
# Paths
NETVM_ROOT = Path("/home/super/Projects/NetVM")
BIN_DIR = NETVM_ROOT / "bin"
JOBS_DIR = NETVM_ROOT / "jobs"
FOLLOWUPS_FILE = NETVM_ROOT / "followups.json"
JOB_LOG = NETVM_ROOT / "job-log.jsonl"
DM_LOG = NETVM_ROOT / "dm-log.jsonl"
DM_PY = BIN_DIR / "dm.py"
DISPATCH_PY = BIN_DIR / "job-dispatch.py"
sys.path.insert(0, str(BIN_DIR))
try:
import pipeline_engine
HAS_PIPELINE = True
except ImportError:
HAS_PIPELINE = False
def utcnow_dt():
return datetime.now(timezone.utc)
def utcnow_str():
return utcnow_dt().isoformat()
def parse_iso(ts_str):
if not ts_str:
return None
try:
ts_clean = ts_str.replace("Z", "+00:00")
dt = datetime.fromisoformat(ts_clean)
if dt.tzinfo is None:
dt = dt.replace(tzinfo=timezone.utc)
return dt
except Exception:
return None
def load_followups():
if not FOLLOWUPS_FILE.exists():
return {}
try:
with open(FOLLOWUPS_FILE, "r", encoding="utf-8") as f:
return json.load(f)
except Exception:
return {}
def save_followups(data):
tmp_path = f"{FOLLOWUPS_FILE}.tmp.{os.getpid()}"
with open(tmp_path, "w", encoding="utf-8") as f:
json.dump(data, f, indent=2)
os.replace(tmp_path, FOLLOWUPS_FILE)
def append_job_log(entry):
os.makedirs(os.path.dirname(os.path.abspath(JOB_LOG)), exist_ok=True)
with open(JOB_LOG, "a", encoding="utf-8") as f:
f.write(json.dumps(entry) + "\n")
def send_dm(sender, recipient, target, text):
"""Dispatch a DM via dm.py. The sweeper's delivery target is explicit
followup state (or explicit in-code escalation routing), so pass the
main-chat opt-in when the target is main (sidechat-first policy)."""
cmd = [
sys.executable,
str(DM_PY),
"send",
"--agent", sender,
"--to", recipient,
"--target", target,
] + (["--allow-main-chat"] if target == "main" else []) + [
text,
]
try:
res = subprocess.run(cmd, capture_output=True, text=True, timeout=90)
return res.returncode == 0, res.stdout.strip() or res.stderr.strip()
except Exception as e:
return False, str(e)
_DM_ID_RE = re.compile(r"\bDM ([0-9a-f]{8})\b")
_THREAD_UUID_RE = re.compile(
r"^[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}$",
re.IGNORECASE,
)
def _resolve_nudge_thread_uuid(nudge_output):
"""Parse the DM id from dm.py's stdout, scan the tail (last 5000 lines)
of dm-log.jsonl for that id's sidechat_autoprovisioned event, falling
back to the sent event's tags.thread. Returns the UUID or None."""
m = _DM_ID_RE.search(nudge_output or "")
if not m:
return None
nudge_id = m.group(1)
fallback = None
try:
with open(DM_LOG, "r", encoding="utf-8") as f:
lines = f.readlines()
except FileNotFoundError:
return None
for line in lines[-5000:]:
line = line.strip()
if not line:
continue
try:
e = json.loads(line)
except Exception:
continue
if e.get("id") != nudge_id:
continue
if e.get("type") == "sidechat_autoprovisioned":
uuid = e.get("thread_uuid") or ""
if _THREAD_UUID_RE.fullmatch(uuid):
return uuid
elif e.get("type") == "sent":
cand = ((e.get("tags") or {}).get("thread")) or ""
if _THREAD_UUID_RE.fullmatch(cand):
fallback = cand
return fallback
_JOB_ID_RE = re.compile(r"^(.+)-(\d{8})-(\d{6})-([0-9a-f]{8})$")
_OP_NAME_RE = re.compile(r"^[a-z][a-z0-9_.]{0,63}$")
_JOB_NAME_RE = re.compile(r"^[a-z0-9-]{1,64}$")
_FALLBACK_RETRY_S = 3600
def fallback_due(rec, now=None):
"""True when a terminal followup should (re)attempt its on_no_result fallback.
Fires once per record; a failed attempt may retry after _FALLBACK_RETRY_S.
Shared by the sweeper terminal path and gravity remediate so the two
firing paths can never double-execute.
"""
fb = rec.get("fallback") or {}
if fb.get("ran"):
return False
ts = fb.get("ts")
if not ts:
return True
try:
last = datetime.fromisoformat(str(ts).replace("Z", "+00:00"))
except Exception:
return True
if last.tzinfo is None:
last = last.replace(tzinfo=timezone.utc)
base = now or utcnow_dt()
return (base - last).total_seconds() >= _FALLBACK_RETRY_S
_EXEC_OPS_MOD = None
def find_job_file(job_name):
"""Locate a job definition in JOBS_DIR or any archive subdirectory."""
p = JOBS_DIR / f"{job_name}.json"
if p.exists():
return p
for match in JOBS_DIR.glob(f"archive/**/{job_name}.json"):
if match.is_file():
return match
for match in JOBS_DIR.glob(f"**/archive/**/{job_name}.json"):
if match.is_file():
return match
return None
def derive_job_name(job_id):
"""Extract the job name from a dispatched job_id (<name>-YYYYMMDD-HHMMSS-<hex8>)."""
m = _JOB_ID_RE.match(job_id or "")
if not m:
return None
name = m.group(1)
if not find_job_file(name):
return None
return name
def load_job_fallback(job_name):
"""Return (spec, error) for a job's on_no_result fallback.
spec is None when the job declares none. Shape:
{"job": "<job-name>"} -> dispatch a fallback job, or
{"op": "<exec-op>", "args": {...}} -> run one exec-constrained op.
"""
job_path = find_job_file(job_name)
if not job_path:
return None, f"unreadable job {job_name}: [Errno 2] No such file or directory: '{JOBS_DIR / (job_name + '.json')}'"
try:
with open(job_path, "r", encoding="utf-8") as f:
cfg = json.load(f)
except Exception as e:
return None, f"unreadable job {job_name}: {e}"
spec = cfg.get("on_no_result")
if spec is None:
return None, None
if not isinstance(spec, dict) or set(spec) - {"job", "op", "args"}:
return None, "on_no_result must be an object with job|op (+args)"
if bool(spec.get("job")) == bool(spec.get("op")):
return None, "on_no_result needs exactly one of job|op"
if spec.get("job"):
jn = spec["job"]
if not isinstance(jn, str) or not _JOB_NAME_RE.fullmatch(jn):
return None, "on_no_result.job must be a valid job name"
if not find_job_file(jn):
return None, f"on_no_result.job {jn!r} does not exist"
else:
if not isinstance(spec.get("op"), str) or not _OP_NAME_RE.fullmatch(spec["op"]):
return None, "on_no_result.op must be a valid op name"
if "args" in spec and not isinstance(spec["args"], dict):
return None, "on_no_result.args must be an object"
return spec, None
def _load_exec_ops():
global _EXEC_OPS_MOD
if _EXEC_OPS_MOD is None:
import importlib.util
mod_spec = importlib.util.spec_from_file_location(
"exec_constrained_sweeper", str(BIN_DIR / "exec-constrained.py"))
mod = importlib.util.module_from_spec(mod_spec)
mod_spec.loader.exec_module(mod)
_EXEC_OPS_MOD = mod
return _EXEC_OPS_MOD
def run_no_result_fallback(rec, dry_run=False):
"""Execute a job's on_no_result fallback at terminal followup expiry.
Returns an outcome dict; never raises (failures are outcome data so
one bad spec can't break the sweep).
"""
dm_id = rec.get("dm_id", "?")
outcome = {"dm_id": dm_id, "ran": False, "mode": None,
"configured": False, "detail": "no job fallback"}
try:
job_name = derive_job_name(rec.get("job_id"))
if not job_name:
outcome["detail"] = "no resolvable job_id"
return outcome
spec, err = load_job_fallback(job_name)
if err:
outcome.update(configured=True, detail=err)
return outcome
if spec is None:
return outcome
outcome["configured"] = True
if dry_run:
outcome.update(mode="dry_run", detail=json.dumps(spec)[:200])
return outcome
if spec.get("job"):
env = os.environ.copy()
env["CHAIN_PREV_JOB_ID"] = rec.get("job_id", "")
env["CHAIN_PREV_RESULT"] = (
f"TIMEOUT: Agent {rec.get('recipient')} gave no result; "
f"on_no_result fallback for job {job_name}")
cmd = [sys.executable, str(DISPATCH_PY), spec["job"]]
try:
p = subprocess.run(cmd, capture_output=True, text=True,
timeout=180, env=env)
except Exception as e:
outcome.update(mode="job", detail=f"dispatch exception: {e}")
return outcome
ok = p.returncode == 0
outcome.update(ran=ok, mode="job",
detail=(f"dispatched {spec['job']}" if ok
else f"dispatch failed: {(p.stderr or p.stdout).strip()[:200]}"))
else:
mod = _load_exec_ops()
op = spec["op"]
op_spec = mod.OPS.get(op)
if op_spec is None:
outcome["detail"] = f"unknown op: {op}"
return outcome
args = dict(spec.get("args") or {})
try:
clean = op_spec["validate"](args)
except Exception as e:
outcome["detail"] = f"op validation failed: {e}"
return outcome
argv = op_spec["build"](clean)
try:
p = subprocess.run(argv, capture_output=True, text=True,
timeout=op_spec.get("timeout", 120))
except Exception as e:
outcome.update(mode="op", detail=f"op exception: {e}")
return outcome
ok = p.returncode == 0
out = (p.stdout or p.stderr or "").strip()
outcome.update(ran=ok, mode="op",
detail=(f"{op} ok: {out[:200]}" if ok
else f"{op} failed rc={p.returncode}: {out[:200]}"))
except Exception as e:
outcome["detail"] = f"fallback exception: {e}"
return outcome
def sweep_cycle(dry_run=False):
followups = load_followups()
if not followups:
return {"status": "ok", "pending": 0, "nudges_sent": 0, "escalations": 0}
now = utcnow_dt()
nudges_count = 0
escalations_count = 0
fallbacks_count = 0
modified = False
for dm_id, rec in list(followups.items()):
if rec.get("status") != "pending":
continue
deadline_dt = parse_iso(rec.get("deadline"))
if not deadline_dt or now < deadline_dt:
continue
# Deadline has expired!
nudges_sent = rec.get("nudges_sent", 0)
nudges_allowed = rec.get("nudges_allowed", 2)
sender = rec.get("sender", "opm")
recipient = rec.get("recipient")
orig_target = rec.get("target", "main")
thread_uuid = rec.get("thread_uuid")
# Ghost-followup fail-fast (2026-10-04, operator-main): a null
# thread_uuid means the sidechat was never provisioned, so the
# harvester can never match a reply. Nudging is pointless -- flag
# for manual triage once instead of burning the nudge budget and
# escalating a ghost. Main-chat followups are unaffected (the
# harvester matches those by target).
if thread_uuid is None and orig_target != "main" and not rec.get("needs_review"):
rec["needs_review"] = True
rec["status"] = "needs_review"
rec["review_reason"] = (
"ghost: unresolvable, manual triage "
f"(thread_uuid null, target={orig_target}; "
"reply can never auto-resolve)"
)
modified = True
append_job_log({
"ts": utcnow_str(),
"type": "followup_ghost_suppressed",
"dm_id": dm_id,
"recipient": recipient,
"target": orig_target,
"reason": "ghost: unresolvable, manual triage",
})
print(f"Sweeper: suppressing ghost followup {dm_id} "
f"(thread_uuid null, target={orig_target}) -> needs_review",
file=sys.stderr)
continue
if nudges_sent < nudges_allowed:
# Deliver next nudge
nudge_num = nudges_sent + 1
is_final = (nudge_num == nudges_allowed)
# Routing: In-thread first, Main on final nudge
delivery_target = "main" if is_final else orig_target
nudge_text = (
f"[nudge {nudge_num}/{nudges_allowed}] [ref:{dm_id}] "
f"Reminder: awaiting reply to request sent at {rec.get('sent_at', 'earlier')}."
)
if is_final and orig_target != "main":
nudge_text += f" (Origin thread: {orig_target})"
print(f"Sweeper: Sending nudge {nudge_num}/{nudges_allowed} to {recipient}/{delivery_target}...")
if not dry_run:
ok, out = send_dm(sender, recipient, delivery_target, nudge_text)
if ok:
nudges_count += 1
rec["nudges_sent"] = nudge_num
rec["last_nudge_at"] = utcnow_str()
# Calculate interval for next nudge: proportional to timeout or default 10m
timeout_s = rec.get("timeout_s", 1800)
step_s = max(300, timeout_s // (nudges_allowed + 1))
rec["deadline"] = (now + timedelta(seconds=step_s)).isoformat()
# C1: backfill the thread the nudge actually landed in.
landed_uuid = _resolve_nudge_thread_uuid(out)
if landed_uuid and rec.get("thread_uuid") != landed_uuid:
rec["thread_uuid"] = landed_uuid
print(f"Sweeper: backfilled thread_uuid={landed_uuid} "
f"for followup {dm_id}", file=sys.stderr)
# C2: record final-nudge routing for the harvester.
if is_final:
rec["final_nudge_target"] = "main"
modified = True
append_job_log({
"ts": utcnow_str(),
"type": "followup_nudged",
"dm_id": dm_id,
"nudge_num": nudge_num,
"recipient": recipient,
"target": delivery_target,
"thread_uuid": rec.get("thread_uuid"),
})
else:
print(f"Sweeper WARNING: nudge send failed: {out}", file=sys.stderr)
# Failure-mode fix (2026-10-04): advance state on send
# failure so a failing nudge is never re-fired every
# timer tick. failed_sends is tracked separately from
# nudges_sent so a delivery failure does not consume a
# real nudge. Backoff: 5 min base, doubling per failure.
failed = rec.get("failed_sends", 0) + 1
rec["failed_sends"] = failed
rec["last_failure_at"] = utcnow_str()
backoff_s = 300 * (2 ** min(failed - 1, 4)) # 5m,10m,20m,40m,80m cap
rec["deadline"] = (now + timedelta(seconds=backoff_s)).isoformat()
modified = True
append_job_log({
"ts": utcnow_str(),
"type": "followup_nudge_failed",
"dm_id": dm_id,
"nudge_num": nudge_num,
"recipient": recipient,
"target": delivery_target,
"failed_sends": failed,
"backoff_s": backoff_s,
"error": str(out)[:200],
})
# After 3 consecutive failures, stop retrying blindly and
# flag for manual review (delivery may be uncertain or the
# target may be permanently broken).
if failed >= 3:
rec["needs_review"] = True
rec["review_reason"] = (
f"nudge send failed {failed} times consecutively "
f"(last: {str(out)[:120]})"
)
append_job_log({
"ts": utcnow_str(),
"type": "followup_needs_review",
"dm_id": dm_id,
"recipient": recipient,
"failed_sends": failed,
})
else:
nudges_count += 1
else:
# All nudges exhausted: Terminal escalation
escalate_to = rec.get("escalate_to", "opm")
esc_text = (
f"[ESCALATION] Agent {recipient} failed to reply to DM {dm_id} "
f"after {nudges_allowed} nudges. Target was: {orig_target} "
f"(thread: {thread_uuid or 'n/a'}). Request sent: {rec.get('sent_at')}."
)
print(f"Sweeper: Escalating expired follow-up {dm_id} to {escalate_to}...")
if not dry_run:
ok, out = send_dm("bl", escalate_to, "main", esc_text)
rec["status"] = "escalated"
rec["escalated_at"] = utcnow_str()
escalations_count += 1
modified = True
append_job_log({
"ts": utcnow_str(),
"type": "followup_escalated",
"dm_id": dm_id,
"recipient": recipient,
"escalated_to": escalate_to,
})
# Check if this dm_id belongs to an active pipeline step
if HAS_PIPELINE:
run_entry, step_entry = pipeline_engine.record_step_timeout(dm_id)
if run_entry and step_entry:
j_name = step_entry.get("job_name")
j_file = JOBS_DIR / f"{j_name}.json"
if j_file.exists():
try:
with open(j_file, "r", encoding="utf-8") as f:
j_cfg = json.load(f)
on_failure = j_cfg.get("on_failure")
if on_failure and (JOBS_DIR / f"{on_failure}.json").exists():
env = os.environ.copy()
env["CHAIN_PREV_JOB_ID"] = step_entry.get("job_id")
env["CHAIN_PREV_RESULT"] = f"TIMEOUT: Agent {recipient} timed out after {nudges_allowed} nudges"
env["CHAIN_PIPELINE_RUN_ID"] = run_entry.get("run_id")
next_step_n = step_entry.get("step_n", 1) + 1
env["CHAIN_STEP_N"] = str(next_step_n)
cmd = [sys.executable, str(DISPATCH_PY), on_failure,
"--pipeline-run", run_entry.get("run_id"),
"--step-n", str(next_step_n)]
subprocess.Popen(cmd, env=env, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
else:
pipeline_engine.fail_pipeline(run_entry.get("run_id"), "step_timed_out_without_fallback")
except Exception:
pass
# on_no_result fallback: the agent never replied, so run
# the job's declared server-side effect now (if any).
# fallback_due() dedupes against the gravity firing path.
fb = (run_no_result_fallback(rec) if fallback_due(rec)
else {"configured": False, "ran": False, "mode": None,
"detail": "fallback already ran"})
if fb["configured"]:
rec["fallback"] = {"ran": fb["ran"], "mode": fb["mode"],
"detail": fb["detail"][:200],
"ts": utcnow_str()}
modified = True
append_job_log({
"ts": utcnow_str(),
"type": ("fallback_executed" if fb["ran"]
else "fallback_failed"),
"dm_id": dm_id,
"recipient": recipient,
"job_id": rec.get("job_id"),
"mode": fb["mode"],
"detail": fb["detail"][:300],
})
print(f"Sweeper: on_no_result fallback for {dm_id}: "
f"ran={fb['ran']} {fb['detail'][:120]}")
if fb["ran"]:
fallbacks_count += 1
else:
escalations_count += 1
if modified and not dry_run:
save_followups(followups)
pending_count = sum(1 for r in followups.values() if r.get("status") == "pending")
return {
"status": "ok",
"pending": pending_count,
"nudges_sent": nudges_count,
"escalations": escalations_count,
"fallbacks": fallbacks_count,
}
def main():
parser = argparse.ArgumentParser(description="Autonomous follow-up deadline tracker and sweeper")
parser.add_argument("--once", action="store_true", help="Run once and exit (default)")
parser.add_argument("--loop", action="store_true", help="Run continuously in a daemon loop")
parser.add_argument("--interval", type=int, default=60, help="Interval in seconds for loop (default 60)")
parser.add_argument("--dry-run", action="store_true", help="Inspect without sending nudges or updating records")
args = parser.parse_args()
if not args.loop:
stats = sweep_cycle(dry_run=args.dry_run)
print(f"[{datetime.now(timezone.utc).strftime('%H:%M:%SZ')}] Sweep cycle: {stats['pending']} pending, {stats['nudges_sent']} nudges, {stats['escalations']} escalations.")
return
print(f"Starting follow-up sweeper loop (interval={args.interval}s)...")
while True:
try:
stats = sweep_cycle(dry_run=args.dry_run)
print(f"[{datetime.now(timezone.utc).strftime('%H:%M:%SZ')}] Sweep cycle: {stats['pending']} pending, {stats['nudges_sent']} nudges, {stats['escalations']} escalations.")
except Exception as e:
print(f"ERROR in sweeper loop: {e}", file=sys.stderr)
time.sleep(args.interval)
if __name__ == "__main__":
main()
-915
View File
@@ -1,915 +0,0 @@
#!/usr/bin/env python3
"""Timer-driven side-chat adoption: shared gravity library.
Pure Python, no network, no bl dependencies — everything the timers need to
decide *where* a message goes and *whether* its loop is alive.
Components:
GravityConfig knobs from gravity.json (fail-closed defaults)
ThreadRegistry (recipient, purpose) -> thread_uuid, JSON-backed
resolve_target() explicit target > registry hit > autocreate > fallback
LoopState FIRING/LANDED/SEND_FAILED/SEEN/ANSWERED/NUDGED/
ESCALATED/CLOSED/BROKEN + transition helpers
detect_breaks() loop-break taxonomy detectors over log-derived events
actionable_digest() wrap a wake digest as a tracked loop message
loop_health() per-agent ANSWERED+CLOSED / LANDED
The live integrations (job-dispatch.py, dm.py, sidechat-wake.py) import the
pure functions here; the patches/ directory shows the call-site diffs.
"""
import json
import os
import re
import sys
import time
from pathlib import Path
# ---------------------------------------------------------------- config
DEFAULTS = {
"job_default_sidechat": True,
"job_sidechat_fallback": "main",
"dm_prefer_sidechat_default": False,
"dm_sidechat_fallback": "main",
"wake_actionable": True,
"wake_ack_timeout": 7200,
"wake_ack_nudges": 1,
"wake_ack_escalate": "opm",
"thread_autocreate": True,
"loop_silent_ticks": 3,
"loop_health_threshold": 0.5,
}
VALID_AGENTS = ("muse", "pip", "646", "opm")
UUID_RE = re.compile(
r"[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}")
class GravityConfig:
"""Knobs. Unknown keys rejected (typo guard); missing keys get
fail-closed defaults (see DEFAULTS)."""
def __init__(self, data=None):
data = data or {}
unknown = set(data) - set(DEFAULTS)
if unknown:
raise ValueError("unknown gravity keys: %s" % sorted(unknown))
self._d = dict(DEFAULTS)
self._d.update(data)
@classmethod
def load(cls, path):
with open(path) as f:
return cls(json.load(f))
def get(self, key):
return self._d[key]
def __getitem__(self, key):
return self._d[key]
def as_dict(self):
return dict(self._d)
# ---------------------------------------------------------------- registry
class ThreadRegistry:
"""(recipient, purpose) -> {thread_uuid, created_ts, last_verified_ts}.
A thread UUID is only valid in the account whose browser created it, so
the registry is keyed by recipient — never share UUIDs across accounts.
Rotation: set()/record_rotation() keep the key pointed at the newest
UUID and stash the old one under `<key>:previous` for audit (same
model as thread_lifecycle.record_rotation).
"""
def __init__(self, path=None, state=None):
self.path = path
self._s = dict(state or {})
@classmethod
def load(cls, path):
try:
with open(path) as f:
return cls(path=path, state=json.load(f))
except (OSError, ValueError):
return cls(path=path)
def save(self):
if not self.path:
raise ValueError("no path for registry save")
tmp = self.path + ".tmp"
with open(tmp, "w") as f:
json.dump(self._s, f, indent=2)
os.replace(tmp, self.path)
def _key(self, recipient, purpose):
return "%s:%s" % (recipient, purpose)
def get(self, recipient, purpose):
"""Current thread UUID for (recipient, purpose), following any
rotation chain. Returns None if unknown."""
return self.resolve(self._key(recipient, purpose))
def resolve(self, key):
"""Return the current UUID for a registry key.
set()/record_rotation() always keep the key pointing at the newest
UUID; the `:previous` entry is audit history, not a traversal chain.
(Mirrors the guarantee in thread_lifecycle.record_rotation.)
"""
cur = self._s.get(key)
if isinstance(cur, str) and UUID_RE.fullmatch(cur):
return cur
return None
def previous(self, recipient, purpose):
"""The rotated-away UUID, if any (for diagnostics)."""
return self._s.get(self._key(recipient, purpose) + ":previous")
def set(self, recipient, purpose, thread_uuid, ts=None):
if not UUID_RE.fullmatch(thread_uuid or ""):
raise ValueError("not a thread UUID: %r" % thread_uuid)
key = self._key(recipient, purpose)
old = self._s.get(key)
if old and old != thread_uuid:
self._s[key + ":previous"] = old
self._s[key] = thread_uuid
if ts:
self._s[key + ":updated_ts"] = ts
def record_rotation(self, recipient, purpose, new_uuid, ts=None):
"""Alias for set() — rotation is just a set with history."""
self.set(recipient, purpose, new_uuid, ts=ts)
def known_purposes(self, recipient):
prefix = recipient + ":"
out = []
for k in self._s:
if k.startswith(prefix) and ":" not in k[len(prefix):]:
out.append(k[len(prefix):])
return sorted(out)
# ---------------------------------------------------------------- target resolution
def resolve_target(recipient, purpose, cfg, registry,
explicit_target=None, probe=None):
"""Decide where a timer-driven send goes.
Order: explicit --target > registry hit (verified) > autocreate >
configured fallback. Never raises for routing reasons — worst case is
the fallback with a logged reason.
probe(thread_uuid) -> bool: cheap reachability check (sidechat use).
create() -> thread_uuid | None: injected by caller (needs browser).
Returns (target, thread_uuid_or_None, reason).
"""
if explicit_target:
tid = UUID_RE.search(explicit_target)
return explicit_target, tid.group(0) if tid else None, "explicit"
if recipient not in VALID_AGENTS:
return cfg["dm_sidechat_fallback"], None, "unknown_recipient"
uuid = registry.get(recipient, purpose)
if uuid:
if probe is None or probe(uuid):
return uuid, uuid, "registry_hit"
return cfg["dm_sidechat_fallback"], None, "registry_stale"
# No mapping: autocreate or fallback.
return None, None, "no_mapping_needs_create"
def fallback_target(cfg):
"""Configured fallback when no sidechat is available. Fails closed
toward visibility: main, never drop."""
return cfg.get("dm_sidechat_fallback") or "main"
# ---------------------------------------------------------------- loop state machine
# Terminal states: CLOSED, ESCALATED, BROKEN, SEND_FAILED.
STATES = ("FIRING", "LANDED", "SEND_FAILED", "SEEN", "ANSWERED",
"NUDGED", "ESCALATED", "CLOSED", "BROKEN")
TERMINAL = frozenset(("SEND_FAILED", "ESCALATED", "CLOSED", "BROKEN"))
# Allowed transitions. NUDGED can cycle (nudge 1..N) until ANSWERED or
# ESCALATED; BROKEN is reachable from any non-terminal state.
TRANSITIONS = {
"FIRING": ("LANDED", "SEND_FAILED", "BROKEN"),
"LANDED": ("SEEN", "ANSWERED", "NUDGED", "BROKEN"),
"SEEN": ("ANSWERED", "NUDGED", "BROKEN"),
"ANSWERED": ("CLOSED", "BROKEN"),
"NUDGED": ("ANSWERED", "NUDGED", "ESCALATED", "BROKEN"),
"SEND_FAILED": (),
"ESCALATED": (),
"CLOSED": (),
"BROKEN": (),
}
class LoopState:
"""One timer-driven loop: timer tick -> tracked follow-up."""
def __init__(self, loop_id, agent, purpose, thread_uuid=None):
self.loop_id = loop_id
self.agent = agent
self.purpose = purpose
self.thread_uuid = thread_uuid
self.state = "FIRING"
self.nudges = 0
self.ticks_without_reply = 0
self.history = [("FIRING", None)]
def transition(self, to, note=None):
if to not in TRANSITIONS.get(self.state, ()):
raise ValueError("illegal loop transition %s -> %s"
% (self.state, to))
self.state = to
if to == "NUDGED":
self.nudges += 1
self.history.append((to, note))
return self
@property
def terminal(self):
return self.state in TERMINAL
@property
def healthy(self):
return self.state in ("ANSWERED", "CLOSED")
def tick(self):
"""One scheduler tick with no reply observed."""
self.ticks_without_reply += 1
return self.ticks_without_reply
# ---------------------------------------------------------------- break detectors
def detect_breaks(loop, cfg, checks):
"""Loop-break taxonomy over caller-supplied check results.
checks: dict with boolean-ish keys:
thread_reachable, signature_ok, scheduler_alive,
reply_observed
Returns a list of break names (may be empty).
"""
breaks = []
if loop.terminal:
return breaks
if not checks.get("scheduler_alive", True):
breaks.append("scheduler_death")
if not checks.get("signature_ok", True):
breaks.append("auth_rot")
if not checks.get("thread_reachable", True):
breaks.append("dead_thread")
# signature rot: thread was rotated — the stored uuid is stale but a
# :previous chain exists. Caller signals via thread_rotated=True and
# should already have resolved; flag if not resolved.
if checks.get("thread_rotated") and not checks.get("thread_resolved"):
breaks.append("signature_rot")
if (not checks.get("reply_observed")
and loop.ticks_without_reply >= cfg["loop_silent_ticks"]
and loop.state in ("LANDED", "SEEN", "NUDGED")):
breaks.append("silent_agent")
return breaks
# ---------------------------------------------------------------- actionable digest
def actionable_digest(digest_text, ack_line=True):
"""Wrap a wake digest so the timer tick becomes a tracked loop.
Adds the ack line that the follow-up record's reply closes. The caller
sends the result via `dm.py send --expect-reply --thread <uuid>` so a
dm_followup record exists — without that send path this is just text.
"""
lines = digest_text.rstrip().split("\n")
if ack_line:
lines.append("")
lines.append("Reply here to acknowledge (closes the loop).")
return "\n".join(lines)
def dm_send_argv(sender, recipient, target_uuid, message, cfg,
purpose="wake"):
"""Build the dm.py argv that makes a timer send a *tracked* loop.
Thread binding (--thread) is what lets nudges route back into the
thread instead of falling back to DM.
"""
return [
"send",
"--agent", sender,
"--to", recipient,
"--target", target_uuid,
"--thread", target_uuid,
"--expect-reply",
"--reply-timeout", str(cfg["wake_ack_timeout"]),
"--reply-nudges", str(cfg["wake_ack_nudges"]),
"--reply-escalate", cfg["wake_ack_escalate"],
"--tag", "purpose:%s" % purpose,
message,
]
# ---------------------------------------------------------------- loop health
def loop_health(loops, threshold=0.5):
"""Per-agent loop health: (ANSWERED + CLOSED) / LANDED.
loops: iterable of LoopState (or dicts with agent/state keys).
Returns {agent: {"landed": n, "answered": n, "health": float|None,
"below_threshold": bool}}.
"""
def _agent(l):
return l.agent if isinstance(l, LoopState) else l.get("agent")
def _state(l):
return l.state if isinstance(l, LoopState) else l.get("state")
# A loop counts as LANDED once it leaves FIRING via LANDED (or beyond).
LANDED_OR_BEYOND = ("LANDED", "SEEN", "ANSWERED", "NUDGED",
"ESCALATED", "CLOSED", "BROKEN")
acc = {}
for l in loops:
a, s = _agent(l), _state(l)
r = acc.setdefault(a, {"landed": 0, "answered": 0})
if s in LANDED_OR_BEYOND:
r["landed"] += 1
if s in ("ANSWERED", "CLOSED"):
r["answered"] += 1
out = {}
for a, r in acc.items():
h = (r["answered"] / r["landed"]) if r["landed"] else None
out[a] = {"landed": r["landed"], "answered": r["answered"],
"health": h,
"below_threshold": h is not None and h < threshold}
return out
# ---------------------------------------------------------------- loop reconstruction & fleet health
NETVM_ROOT = "/home/super/Projects/NetVM"
FOLLOWUPS_FILE = os.path.join(NETVM_ROOT, "followups.json")
DM_LOG_FILE = os.path.join(NETVM_ROOT, "dm-log.jsonl")
JOB_LOG_FILE = os.path.join(NETVM_ROOT, "job-log.jsonl")
def reconstruct_loops(limit=50, agent=None, status_filter=None) -> list:
"""Reconstruct active and recent loops from followups.json and dm-log.jsonl.
Returns a list of dicts:
loop_id, agent, sender, target, purpose, state, sent_at, deadline,
nudges_sent, nudges_allowed, escalate_to, tags, summary
"""
loops = {} # loop_id -> dict
# 1. Load active/persisted followups
if os.path.exists(FOLLOWUPS_FILE):
try:
with open(FOLLOWUPS_FILE, "r") as f:
fdata = json.load(f)
if isinstance(fdata, dict):
for did, rec in fdata.items():
recipient = rec.get("recipient")
st = rec.get("status", "pending")
# Map status to loop state
if st == "pending":
lstate = "NUDGED" if rec.get("nudges_sent", 0) > 0 else "LANDED"
elif st == "escalated":
lstate = "ESCALATED"
elif st in ("resolved", "closed"):
lstate = "CLOSED"
else:
lstate = st.upper()
loops[did] = {
"loop_id": did,
"agent": recipient,
"sender": rec.get("sender", "super"),
"target": rec.get("target", "main"),
"thread_uuid": rec.get("thread_uuid"),
"purpose": rec.get("route") or "followup",
"state": lstate,
"sent_at": rec.get("sent_at", ""),
"deadline": rec.get("deadline", ""),
"timeout_s": rec.get("timeout_s", 3600),
"nudges_sent": rec.get("nudges_sent", 0),
"nudges_allowed": rec.get("nudges_allowed", 2),
"escalate_to": rec.get("escalate_to", "opm"),
"tags": {"reply:expected": True},
"summary": f"Follow-up for DM {did}",
"source": "followups.json",
}
except Exception:
pass
# 2. Extract tracked DMs from dm-log.jsonl
answers_map = {} # id or ref -> answer event
dm_events = []
failed_ids = set() # DM ids that failed to send - never delivered
if os.path.exists(DM_LOG_FILE):
try:
with open(DM_LOG_FILE, "r") as f:
for line in f:
line = line.strip()
if not line:
continue
try:
entry = json.loads(line)
ev_type = entry.get("type", "")
msg = entry.get("msg") or ""
# Check for answers/results/acks
# Note: "verified" means delivered, NOT answered - do not include it here
if "[RESULT" in msg or "[ACK" in msg or "acknowledged" in msg:
sender = entry.get("agent", "")
# Extract referenced DM id if present
m_ref = re.search(r"\[ref:([a-f0-9-]+)\]", msg) or re.search(r"\[(?:ACK|RESULT)\s+([a-f0-9-]+)", msg)
if m_ref:
answers_map[m_ref.group(1)] = entry
if entry.get("id"):
answers_map[entry["id"]] = entry
# Track failed sends - these were never delivered
if ev_type in ("send_failed", "failed"):
if entry.get("id"):
failed_ids.add(entry.get("id"))
tags = entry.get("tags") or {}
if tags.get("reply:expected") or "[reply:expected]" in msg:
dm_events.append(entry)
except Exception:
continue
except Exception:
pass
# Merge dm_events into loops
for entry in dm_events:
did = entry.get("id") or entry.get("msg_id") or ""
if not did:
continue
if did in loops:
continue # Already have active record from followups.json
if did in failed_ids:
continue # Send failed - never delivered, don't count against agent health
tags = entry.get("tags") or {}
recipient = entry.get("to") or entry.get("recipient") or ""
sender = entry.get("agent") or entry.get("sender") or "super"
target = entry.get("target") or "main"
sent_at = entry.get("ts") or entry.get("sent_at") or ""
timeout_s = int(tags.get("reply:timeout") or 3600)
nudges_max = int(tags.get("reply:nudges") or 2)
esc = tags.get("reply:escalate") or "opm"
purpose = tags.get("route") or "dm"
# Determine state
if did in answers_map or f"ref:{did}" in answers_map:
lstate = "CLOSED"
else:
# Check age
lstate = "LANDED"
loops[did] = {
"loop_id": did,
"agent": recipient,
"sender": sender,
"target": target,
"thread_uuid": tags.get("thread"),
"purpose": purpose,
"state": lstate,
"sent_at": sent_at,
"deadline": "",
"timeout_s": timeout_s,
"nudges_sent": 0,
"nudges_allowed": nudges_max,
"escalate_to": esc,
"tags": tags,
"summary": entry.get("msg", "")[:80],
"source": "dm-log.jsonl",
}
# Filter and sort
result = list(loops.values())
if agent:
result = [l for l in result if l.get("agent") == agent or l.get("sender") == agent]
if status_filter:
sf = status_filter.lower()
if sf == "active" or sf == "pending":
result = [l for l in result if l.get("state") in ("FIRING", "LANDED", "SEEN", "NUDGED", "ESCALATED")]
elif sf in ("closed", "resolved"):
result = [l for l in result if l.get("state") in ("CLOSED", "ANSWERED")]
else:
result = [l for l in result if l.get("state", "").lower() == sf]
# Sort newest first
result.sort(key=lambda x: x.get("sent_at", ""), reverse=True)
return result[:limit]
def get_fleet_loop_health(threshold=None) -> dict:
"""Calculate fleet loop health per agent and overall verdict."""
if threshold is None:
try:
from variables import Variables
threshold = Variables().get_float("loop_health_threshold")
except Exception:
threshold = 0.5
loops = reconstruct_loops(limit=100)
raw = loop_health(loops, threshold=threshold)
out = {}
for ag in VALID_AGENTS:
info = raw.get(ag, {"landed": 0, "answered": 0, "health": None, "below_threshold": False})
landed = info.get("landed", 0)
answered = info.get("answered", 0)
h = info.get("health")
if landed == 0:
status = "IDLE"
elif h is not None and h >= threshold:
status = "HEALTHY"
else:
status = "DEGRADED"
out[ag] = {
"agent": ag,
"landed": landed,
"answered": answered,
"health": h,
"health_pct": f"{int(h * 100)}%" if h is not None else "-",
"threshold": threshold,
"below_threshold": info.get("below_threshold", False),
"status": status,
}
total_landed = sum(v["landed"] for v in out.values())
total_answered = sum(v["answered"] for v in out.values())
overall_h = (total_answered / total_landed) if total_landed > 0 else None
return {
"agents": out,
"summary": {
"total_landed": total_landed,
"total_answered": total_answered,
"overall_health": overall_h,
"overall_health_pct": f"{int(overall_h * 100)}%" if overall_h is not None else "-",
"threshold": threshold,
"healthy": overall_h is None or overall_h >= threshold,
}
}
def diagnose_breaks() -> list:
"""Diagnose break taxonomy across intrinsic loops and support services."""
import subprocess
breaks = []
# 1. Scheduler / Timers check
try:
r = subprocess.run(["systemctl", "--user", "is-system-running"],
capture_output=True, text=True, timeout=5)
sys_state = r.stdout.strip()
if sys_state in ("offline", "stopped"):
breaks.append({
"type": "scheduler_death",
"severity": "CRITICAL",
"component": "systemd",
"detail": f"Systemd user instance is {sys_state}",
"remedy": "Restart systemd user session or start timer jobs manually."
})
except Exception as e:
breaks.append({
"type": "scheduler_death",
"severity": "WARNING",
"component": "systemd",
"detail": f"Could not check systemd status: {e}",
"remedy": "Verify systemctl --user is available."
})
# 2. SSH key signature check
key_path = os.path.expanduser("~/.ssh/id_ed25519")
if not os.path.exists(key_path):
breaks.append({
"type": "auth_rot",
"severity": "CRITICAL",
"component": "ssh-keys",
"detail": f"SSH signing key {key_path} not found",
"remedy": "Generate ed25519 key at ~/.ssh/id_ed25519 for cryptographically signed DMs."
})
# 3. Active follow-up loops check
active_loops = reconstruct_loops(limit=20, status_filter="pending")
now_ts = time.time()
for l in active_loops:
nudges_sent = l.get("nudges_sent", 0)
nudges_max = l.get("nudges_allowed", 2)
if nudges_sent >= nudges_max and l.get("state") == "ESCALATED":
breaks.append({
"type": "silent_agent",
"severity": "WARNING",
"loop_id": l["loop_id"],
"agent": l["agent"],
"detail": f"Agent {l['agent']} silent after {nudges_sent}/{nudges_max} nudges for DM {l['loop_id']}",
"remedy": f"Check agent {l['agent']} browser tab with 'super fleet status' or nudge via 'super dm send'."
})
# 4. Check for agents held up on approvals
try:
import approvals
fleet_apps = approvals.check_fleet_approvals()
for app in fleet_apps:
if app.get("has_pending"):
node = app["node"]
is_trusted = app.get("is_trusted", False)
ip = app.get("target") or app.get("ip") or "unknown target"
breaks.append({
"type": "approval_blocked",
"severity": "WARNING" if is_trusted else "CRITICAL",
"component": f"node:{node}",
"agent": node,
"detail": f"Agent {node} is held up on browser approval for {ip}",
"remedy": f"Run 'box approvals auto' or 'box approvals allow {node}'."
})
for w in app.get("input_waits") or []:
node = app["node"]
breaks.append({
"type": "input_wait",
"severity": "WARNING",
"component": f"node:{node}",
"agent": node,
"detail": f"Agent {node} task '{w.get('task')}' is waiting: {w.get('status')} ({w.get('when')})",
"remedy": f"Open {node}'s task and answer it, or 'box approvals check --node {node}'."
})
except Exception:
pass
return breaks
def _load_sweeper_module():
import importlib.util
mod_spec = importlib.util.spec_from_file_location(
"followup_sweeper_gravity",
str(Path(__file__).resolve().parent / "followup-sweeper.py"))
mod = importlib.util.module_from_spec(mod_spec)
mod_spec.loader.exec_module(mod)
return mod
def maybe_run_terminal_fallback(rec, now_iso, dry_run=False):
"""Run a job's on_no_result fallback once at terminal followup expiry.
Returns an outcome dict, or None when the record has no resolvable
fallback. Never raises. The sweeper terminal path shares the
fallback_due() guard, so the two firing paths can't double-execute.
"""
try:
sw = _load_sweeper_module()
except Exception as e:
return {"ran": False, "mode": None,
"detail": f"sweeper import failed: {e}"}
try:
if not sw.fallback_due(rec):
return None
job_name = sw.derive_job_name(rec.get("job_id"))
if not job_name:
return None
spec, err = sw.load_job_fallback(job_name)
if err or spec is None:
return None
if dry_run:
return {"ran": False, "mode": "dry_run",
"detail": json.dumps(spec)[:200]}
out = sw.run_no_result_fallback(rec)
rec["fallback"] = {"ran": out["ran"], "mode": out["mode"],
"detail": out["detail"][:200], "ts": now_iso}
try:
sw.append_job_log({
"ts": now_iso,
"type": ("fallback_executed" if out["ran"]
else "fallback_failed"),
"dm_id": rec.get("dm_id"),
"recipient": rec.get("recipient"),
"job_id": rec.get("job_id"),
"mode": out["mode"],
"detail": out["detail"][:300],
})
except Exception:
pass
return out
except Exception as e:
return {"ran": False, "mode": None,
"detail": f"fallback exception: {e}"}
def remediate_breaks(dry_run=False) -> dict:
"""Progressively auto-remediate soft loop breakages while escalating hard breakages.
Soft breakages (auto-healed):
- Pending followups that have received an answer in dm-log.jsonl or chat-history
are resolved.
- Pending followups with expired deadlines and nudges remaining are re-armed
and swept immediately.
Hard breakages (escalated loudly):
- silent_agent (nudges exhausted, no response)
- auth_rot (missing SSH signing keys)
- scheduler_death (systemd user session offline)
"""
import subprocess
from datetime import datetime, timezone
remediated = []
escalated = []
# 1. Check diagnosed hard breaks first
breaks = diagnose_breaks()
for b in breaks:
if b.get("severity") in ("CRITICAL", "WARNING"):
escalated.append(b)
# 2. Check followups.json for soft break healing
f_path = Path("/home/super/Projects/NetVM/followups.json")
f_modified = False
rearm_sweeper = False
if f_path.exists():
try:
with open(f_path, "r") as f:
fdata = json.load(f)
except Exception:
fdata = {}
now_iso = datetime.now(timezone.utc).isoformat()
# Build answer map from reconstruct_loops
loops = reconstruct_loops(limit=200)
answered_dms = {
l["loop_id"]: l for l in loops if l.get("state") in ("ANSWERED", "CLOSED")
}
for dm_id, rec in fdata.items():
if rec.get("status") == "pending":
# Check if it was actually answered in logs
if dm_id in answered_dms:
remediated.append({
"action": "auto_resolve_answered",
"loop_id": dm_id,
"agent": rec.get("recipient"),
"detail": f"Follow-up {dm_id} received reply in log but was pending in followups.json. Marked resolved."
})
if not dry_run:
rec["status"] = "resolved"
rec["resolved_at"] = now_iso
rec["resolved_note"] = "auto-healed: reply detected in dm-log"
f_modified = True
continue
# Check if deadline expired and nudges remaining
dl_str = rec.get("deadline", "")
nudges_sent = rec.get("nudges_sent", 0)
nudges_allowed = rec.get("nudges_allowed", 2)
is_expired = False
if dl_str:
try:
dl_dt = datetime.fromisoformat(dl_str.replace("Z", "+00:00"))
if datetime.now(timezone.utc) > dl_dt:
is_expired = True
except Exception:
pass
if is_expired and nudges_sent < nudges_allowed:
remediated.append({
"action": "rearm_expired_nudge",
"loop_id": dm_id,
"agent": rec.get("recipient"),
"detail": f"Deadline expired for {dm_id} ({nudges_sent}/{nudges_allowed} nudges). Re-arming immediate sweep."
})
if not dry_run:
rec["deadline"] = now_iso
f_modified = True
rearm_sweeper = True
# Terminal: nudges exhausted and still no reply. Run the
# job's on_no_result fallback (server-side guarantee).
if is_expired and nudges_sent >= nudges_allowed:
fb = maybe_run_terminal_fallback(rec, now_iso, dry_run)
if fb is not None:
remediated.append({
"action": "terminal_fallback",
"loop_id": dm_id,
"agent": rec.get("recipient"),
"detail": f"on_no_result ran={fb.get('ran')}: {fb.get('detail', '')[:160]}",
})
if not dry_run and "fallback" in rec:
f_modified = True
if f_modified and not dry_run:
tmp = f"{f_path}.tmp.{os.getpid()}"
with open(tmp, "w") as f:
json.dump(fdata, f, indent=2)
os.replace(tmp, f_path)
if rearm_sweeper and not dry_run:
sweeper_py = Path("/home/super/Projects/NetVM/bin/followup-sweeper.py")
if sweeper_py.exists():
try:
# A single nudge send takes ~10s median; give the sweep
# room to finish or it dies mid-first-send every time.
subprocess.run([sys.executable, str(sweeper_py), "--once"],
timeout=300)
except Exception:
pass
# Auto-remediate trusted approval blocks
try:
import approvals
fleet_apps = approvals.check_fleet_approvals()
for app in fleet_apps:
if app.get("has_pending") and app.get("is_trusted") and app.get("status") != "KEY_APPROVAL":
node = app["node"]
if not dry_run:
approvals.allow_node_approval(node, always=True, caller="loop-remediate")
remediated.append({
"type": "approval_auto_allowed",
"agent": node,
"target": app.get("ip"),
"action": f"Auto-approved trusted browser request on {node} ({app.get('ip')})"
})
except Exception:
pass
# 3. Alert on hard breakages if any exist
if escalated and not dry_run:
job_log = Path("/home/super/Projects/NetVM/job-log.jsonl")
alert_msg = f"[HARD_BREAK_ALERT] Detected {len(escalated)} unresolvable loop failure(s): " + "; ".join(
f"{b.get('type')} ({b.get('severity')}): {b.get('detail')}" for b in escalated[:3]
)
try:
with open(job_log, "a", encoding="utf-8") as jf:
jf.write(json.dumps({
"ts": datetime.now(timezone.utc).isoformat(),
"type": "hard_break_alert",
"escalated_count": len(escalated),
"items": escalated,
"summary": alert_msg[:280]
}) + "\n")
except Exception:
pass
# Dispatch DM alert to opm
dm_py = Path("/home/super/Projects/NetVM/bin/dm.py")
if dm_py.exists():
try:
subprocess.run([
sys.executable, str(dm_py), "send",
"--agent", "super",
"--to", "opm",
"--target", "main",
alert_msg[:800]
], timeout=15, capture_output=True)
except Exception:
pass
return {
"ok": True,
"remediated": remediated,
"escalated": escalated,
"dry_run": dry_run,
"count": len(remediated),
}
def main(argv=None):
import argparse
ap = argparse.ArgumentParser(description="Loop gravity: reconcile and remediate followup loops")
ap.add_argument("--remediate", action="store_true",
help="Run remediate_breaks once (what loop-remediator.timer invokes)")
ap.add_argument("--dry-run", action="store_true",
help="Report actions without writing state or sending anything")
args = ap.parse_args(argv)
if not args.remediate:
ap.print_help()
return 2
result = remediate_breaks(dry_run=args.dry_run)
print(json.dumps(result, indent=2))
return 0
if __name__ == "__main__":
sys.exit(main())
-53
View File
@@ -1,53 +0,0 @@
"""hatch_menu — modular muse.ai settings-menu navigation + toggles.
One module per part so a site change means patching one file:
mouse.py trusted input primitives (real click, escape, close)
dialog.py settings dialog open / tab nav / rows / back / text
controls.py generic radio/switch primitives (verify-then-fallback)
toggles.py toggle registry + sessions over controls and tabs
tabs/ one module per settings tab (uniform describe())
invite.py consumes this package; `box chromebox` exposes toggles.
"""
from hatch_menu.dialog import (
TAB_NAMES,
TABS,
click_row,
close_settings,
describe_rows,
dialog_present,
dialog_text,
go_back,
goto_tab,
open_settings,
)
from hatch_menu.mouse import MouseError, close, escape, real_click
from hatch_menu.toggles import (
MenuError,
describe_tab,
get_toggle,
list_toggles,
set_toggle,
)
__all__ = [
"TAB_NAMES",
"TABS",
"MenuError",
"MouseError",
"click_row",
"close",
"close_settings",
"describe_rows",
"describe_tab",
"dialog_present",
"dialog_text",
"escape",
"get_toggle",
"go_back",
"goto_tab",
"list_toggles",
"open_settings",
"real_click",
]
-248
View File
@@ -1,248 +0,0 @@
"""Generic control primitives (ws-level, no tab knowledge).
Radio/switch list + set with verify-then-fallback: synthetic click,
verify state flipped, else trusted real click, verify again. Tab
modules build their flows on these; toggles.py adds addressing.
"""
import time
from approvals import cdp_evaluate
from hatch_menu.mouse import real_click
JS_LIST_RADIOS = """(() => {
const d = document.querySelector('[role="dialog"]');
if (!d) return null;
return Array.from(d.querySelectorAll('input[type="radio"]')).map(r => {
let head = '', el = r.parentElement, depth = 0;
while (el && el !== d && depth < 6) {
const h = el.querySelector('h1,h2,h3,h4');
if (h && (h.innerText || '').trim()) {
head = h.innerText.trim().slice(0, 60);
break;
}
el = el.parentElement;
depth += 1;
}
const lab = r.closest('label');
return {value: r.value, checked: !!r.checked, heading: head,
aria: r.getAttribute('aria-label') || '',
label: lab ? (lab.innerText || '').trim().slice(0, 80) : ''};
});
})()"""
JS_CLICK_RADIO = """((heading, value) => {
const d = document.querySelector('[role="dialog"]');
if (!d) return 'NO_DIALOG';
const radios = Array.from(d.querySelectorAll('input[type="radio"]'));
const headOf = (r) => {
let el = r.parentElement, depth = 0;
while (el && el !== d && depth < 6) {
const h = el.querySelector('h1,h2,h3,h4');
if (h && (h.innerText || '').trim())
return h.innerText.trim().toLowerCase();
el = el.parentElement;
depth += 1;
}
return '';
};
const t = radios.find(r => headOf(r) === heading.toLowerCase()
&& r.value === value);
if (!t) return 'NO_MATCH';
t.click();
return 'CLICKED';
})('%s', '%s')"""
JS_RADIO_RECT = """((heading, value) => {
const d = document.querySelector('[role="dialog"]');
if (!d) return null;
const radios = Array.from(d.querySelectorAll('input[type="radio"]'));
const headOf = (r) => {
let el = r.parentElement, depth = 0;
while (el && el !== d && depth < 6) {
const h = el.querySelector('h1,h2,h3,h4');
if (h && (h.innerText || '').trim())
return h.innerText.trim().toLowerCase();
el = el.parentElement;
depth += 1;
}
return '';
};
const t = radios.find(r => headOf(r) === heading.toLowerCase()
&& r.value === value);
if (!t) return null;
const r = t.getBoundingClientRect();
return {x: r.x + r.width / 2, y: r.y + r.height / 2};
})('%s', '%s')"""
JS_CLICK_RADIO_ARIA = """((name) => {
const d = document.querySelector('[role="dialog"]');
if (!d) return 'NO_DIALOG';
const n = name.toLowerCase();
const t = Array.from(d.querySelectorAll('input[type="radio"]'))
.find(r => (r.getAttribute('aria-label') || '').toLowerCase() === n
|| (r.value || '').toLowerCase() === n);
if (!t) return 'NO_MATCH';
t.click();
return 'CLICKED';
})('%s')"""
JS_RADIO_ARIA_RECT = """((name) => {
const d = document.querySelector('[role="dialog"]');
if (!d) return null;
const n = name.toLowerCase();
const t = Array.from(d.querySelectorAll('input[type="radio"]'))
.find(r => (r.getAttribute('aria-label') || '').toLowerCase() === n
|| (r.value || '').toLowerCase() === n);
if (!t) return null;
const r = t.getBoundingClientRect();
return {x: r.x + r.width / 2, y: r.y + r.height / 2};
})('%s')"""
JS_LIST_SWITCHES = """(() => {
const d = document.querySelector('[role="dialog"]');
if (!d) return null;
return Array.from(d.querySelectorAll('[role="switch"]')).map(s => {
let el = s.parentElement, label = '', depth = 0;
while (el && el !== d && depth < 6) {
const t = (el.innerText || '').trim().replace(/\\s+/g, ' ');
if (t && t.length < 250) { label = t; break; }
el = el.parentElement;
depth += 1;
}
return {label: label.slice(0, 200),
aria: s.getAttribute('aria-label') || '',
checked: s.getAttribute('aria-checked') === 'true'};
});
})()"""
JS_CLICK_SWITCH = """((label) => {
const d = document.querySelector('[role="dialog"]');
if (!d) return 'NO_DIALOG';
const rowOf = (s) => {
let el = s.parentElement, depth = 0;
while (el && el !== d && depth < 6) {
const t = (el.innerText || '').trim();
if (t && t.length < 250) return t.toLowerCase();
el = el.parentElement;
depth += 1;
}
return '';
};
const hits = Array.from(d.querySelectorAll('[role="switch"]'))
.filter(s => rowOf(s).includes(label.toLowerCase())
|| (s.getAttribute('aria-label') || '').toLowerCase()
.includes(label.toLowerCase()));
if (!hits.length) return 'NO_MATCH';
if (hits.length > 1) return 'AMBIGUOUS';
hits[0].click();
return 'CLICKED';
})('%s')"""
def _eval(ws, js, timeout=8.0):
try:
return cdp_evaluate(ws, js, timeout=timeout)
except Exception:
return None
VERIFY_TRIES = 10
VERIFY_PAUSE = 1.5
def _poll(check, tries=VERIFY_TRIES, pause=VERIFY_PAUSE):
"""Poll a state check until true. Fast exit; tolerates slow commits."""
for _ in range(tries):
try:
if check():
return True
except Exception:
pass
time.sleep(pause)
return False
def list_radios(ws):
"""All dialog radios with heading/label/value/checked (or [])."""
rows = _eval(ws, JS_LIST_RADIOS)
return rows if isinstance(rows, list) else []
def list_switches(ws):
"""All dialog switches with row label + checked (or [])."""
rows = _eval(ws, JS_LIST_SWITCHES)
return rows if isinstance(rows, list) else []
def radio_state(ws, heading, value):
"""Checked state of one heading-grouped radio, None if absent."""
for r in list_radios(ws):
if (r.get("heading") or "").lower() == heading.lower() \
and r.get("value") == value:
return bool(r.get("checked"))
return None
def set_radio_by_heading(ws, heading, value):
"""Set a heading-grouped radio; poll, else trusted click, poll."""
if _eval(ws, JS_CLICK_RADIO % (heading, value)) == "CLICKED" \
and _poll(lambda: radio_state(ws, heading, value) is True):
return True
rect = _eval(ws, JS_RADIO_RECT % (heading, value))
if not rect or "x" not in rect:
return False
try:
real_click(ws, rect["x"], rect["y"])
except Exception:
return False
return _poll(lambda: radio_state(ws, heading, value) is True)
def radio_aria_state(ws, name):
"""Checked state of one aria-labeled radio, None if absent."""
n = name.lower()
for r in list_radios(ws):
if (r.get("aria") or "").lower() == n \
or (r.get("value") or "").lower() == n:
return bool(r.get("checked"))
return None
def set_radio_by_aria(ws, name):
"""Set an aria-labeled radio; poll, else trusted click, poll."""
if _eval(ws, JS_CLICK_RADIO_ARIA % name) == "CLICKED" \
and _poll(lambda: radio_aria_state(ws, name) is True):
return True
rect = _eval(ws, JS_RADIO_ARIA_RECT % name)
if not rect or "x" not in rect:
return False
try:
real_click(ws, rect["x"], rect["y"])
except Exception:
return False
return _poll(lambda: radio_aria_state(ws, name) is True)
def switch_state(ws, label):
"""Checked state of one label-matched switch, None if not unique."""
matches = [s for s in list_switches(ws)
if label.lower() in (s.get("label") or "").lower()
or label.lower() in (s.get("aria") or "").lower()]
if len(matches) != 1:
return None
return bool(matches[0].get("checked"))
def set_switch(ws, label, on):
"""Set a switch by row-label/aria match; no-op when already there.
Single click only: a fallback re-click would UNDO a slow commit
(switches toggle). The fresh-session readback decides.
"""
state = switch_state(ws, label)
if state is None:
return False
if state == bool(on):
return True
if _eval(ws, JS_CLICK_SWITCH % label) != "CLICKED":
return False
return _poll(lambda: switch_state(ws, label) is bool(on))
-162
View File
@@ -1,162 +0,0 @@
"""Settings dialog navigation: open, tabs, rows, back, text.
Open is retried (single-shot opens flake ~1/4 live): each attempt
re-Escapes and re-drives the dock menu from scratch. Tab clicks are
idempotent (re-clicking the active tab is a harmless no-op), so goto
always clicks and reports the click result instead of guessing which
tab is active.
"""
import json
import time
from approvals import cdp_evaluate
from hatch_menu.mouse import escape, real_click
TAB_NAMES = ["General", "Connectors", "Wallet", "Secure store",
"Permissions", "Messaging channels", "Devices",
"Data controls", "Help & support", "Legal info"]
TABS = TAB_NAMES # legacy alias
# Tab rail buttons carry bare tab names and live outside any nav
# landmark, so row matchers exclude them by exact text (live 2026-10-06).
_JS_TABS = json.dumps(TAB_NAMES)
DOCK_MORE_TESTID = "hatch-dock-more"
JS_DOCK_RECT = ("(() => { const b = document.querySelector("
"'[data-testid=\"hatch-dock-more\"]'); if (!b) return null;"
" const r = b.getBoundingClientRect();"
" return {x: r.x + r.width/2, y: r.y + r.height/2}; })()")
JS_CLICK_SETTINGS_ITEM = ("(() => { const it = Array.from(document."
"querySelectorAll('[role=\"menuitem\"]')).find(el => el.getAttribute("
"'data-pel-click') === 'settings_nav_click');"
" if (!it) return 'NO_ITEM'; it.click(); return 'CLICKED'; })()")
JS_DIALOG_PRESENT = ("(() => !!document.querySelector('[role=\"dialog\"]'))()")
JS_DIALOG_TEXT = ("(() => { const d = document.querySelector("
"'[role=\"dialog\"]'); return d ? d.innerText : null; })()")
JS_GOTO_TAB_TMPL = ("(() => { const b = Array.from(document."
"querySelectorAll('[role=\"dialog\"] button')).find(x => "
"(x.innerText||'').trim() === '%s');"
" if (!b) return 'NO_TAB'; b.click(); return 'CLICKED'; })()")
JS_CLICK_ROW_TMPL = ("((name) => {"
" const d = document.querySelector('[role=\"dialog\"]');"
" if (!d) return 'NO_DIALOG';"
" const TABS = " + _JS_TABS + ";"
" const inNav = (el) => !!el.closest("
"'nav, [role=\"tablist\"], [role=\"navigation\"]');"
" const els = Array.from(d.querySelectorAll("
"'button, [role=\"button\"], a')).filter(e => !inNav(e));"
" const t = els.find(e => {"
" const txt = (e.innerText || '').trim();"
" return !TABS.includes(txt) && txt.toLowerCase()"
".startsWith(name.toLowerCase()); });"
" if (!t) return 'NO_ROW'; t.click(); return 'CLICKED'; })('%s')")
JS_DESCRIBE_ROWS = ("(() => {"
" const d = document.querySelector('[role=\"dialog\"]');"
" if (!d) return null;"
" const TABS = " + _JS_TABS + ";"
" const inNav = (el) => !!el.closest("
"'nav, [role=\"tablist\"], [role=\"navigation\"]');"
" return Array.from(d.querySelectorAll("
"'button, [role=\"button\"], a')).filter(e => !inNav(e))"
".filter(e => !TABS.includes((e.innerText || '').trim()))"
".map(e => { const lines = (e.innerText || '').trim().split('\\n');"
" return {name: (lines[0] || '').slice(0, 80),"
" detail: lines.slice(1).join(' / ').slice(0, 120)}; }); })()")
JS_GO_BACK = ("(() => { const d = document.querySelector("
"'[role=\"dialog\"]'); if (!d) return 'NO_DIALOG';"
" const b = Array.from(d.querySelectorAll('button')).find("
"x => (x.getAttribute('aria-label') || '') === 'Go back');"
" if (!b) return 'NO_BACK'; b.click(); return 'CLICKED'; })()")
def _eval(ws, js, timeout=8.0):
try:
return cdp_evaluate(ws, js, timeout=timeout)
except Exception:
return None
def dialog_present(ws):
"""True when a dialog is currently open."""
return bool(_eval(ws, JS_DIALOG_PRESENT, timeout=5.0))
def dialog_text(ws, timeout=8.0, limit=4000):
"""Inner text of the open dialog, or None."""
try:
text = cdp_evaluate(ws, JS_DIALOG_TEXT, timeout=timeout)
except Exception:
return None
if not isinstance(text, str) or not text:
return None
return text[:limit]
def open_settings(ws, tries=3):
"""Open the Settings dialog via the dock menu. True when open."""
for _ in range(tries):
escape(ws)
time.sleep(0.5)
rect = _eval(ws, JS_DOCK_RECT, timeout=5.0)
if not rect or "x" not in rect:
continue
try:
real_click(ws, rect["x"], rect["y"])
except Exception:
continue
time.sleep(1.0)
if _eval(ws, JS_CLICK_SETTINGS_ITEM,
timeout=5.0) != "CLICKED":
continue
time.sleep(1.5)
if dialog_present(ws):
return True
return False
def goto_tab(ws, name, timeout=5.0):
"""Click a Settings tab by visible name. True when clicked."""
if _eval(ws, JS_GOTO_TAB_TMPL % name, timeout=timeout) != "CLICKED":
return False
time.sleep(0.8)
return True
def click_row(ws, name, tab=None, timeout=8.0):
"""Click a content row; verify the drill/expand opened. Bool."""
if tab is not None and not goto_tab(ws, tab, timeout=timeout):
return False
if _eval(ws, JS_CLICK_ROW_TMPL % name, timeout=timeout) != "CLICKED":
return False
time.sleep(1.0)
text = dialog_text(ws, timeout=timeout) or ""
return name.lower() in text.lower()
def describe_rows(ws, tab=None, timeout=8.0):
"""Inventory rows (name/detail) on a tab. [] when unreadable."""
if tab is not None and not goto_tab(ws, tab, timeout=timeout):
return []
rows = _eval(ws, JS_DESCRIBE_ROWS, timeout=timeout)
return rows if isinstance(rows, list) else []
def go_back(ws):
"""Click the sub-page Go back button. True when clicked."""
if _eval(ws, JS_GO_BACK, timeout=5.0) != "CLICKED":
return False
time.sleep(0.8)
return True
def close_settings(ws):
"""Dismiss settings/popovers. Never raises."""
escape(ws)
-62
View File
@@ -1,62 +0,0 @@
"""Trusted input primitives for menu automation.
Radix triggers (dock menu, permission-mode choosers) ignore synthetic
JS clicks: they need real CDP Input.dispatchMouseEvent press+release.
"""
import json
import time
from approvals import _cdp_req_ids as _shared_cdp_ids, cdp_evaluate
ESCAPE_JS = ("(() => { document.dispatchEvent(new KeyboardEvent("
"'keydown', {key: 'Escape', code: 'Escape',"
" bubbles: true})); return 'ESC'; })()")
class MouseError(RuntimeError):
"""Trusted click failed (CDP transport or echo timeout)."""
def escape(ws):
"""Dismiss topmost popover/menu/dialog. Never raises."""
try:
cdp_evaluate(ws, ESCAPE_JS, timeout=3.0)
except Exception:
pass
def close(ws):
"""Close a CDP websocket. Never raises."""
try:
ws.close()
except Exception:
pass
def real_click(ws, x, y, timeout=5.0):
"""Trusted press+release at page coordinates.
Shares approvals' monotonic CDP id counter so ids stay unique on
the connection; matches responses by id like cdp_evaluate.
Raises MouseError on transport or echo-timeout failure.
"""
try:
for typ in ("mousePressed", "mouseReleased"):
req_id = next(_shared_cdp_ids)
ws.send(json.dumps({"id": req_id,
"method": "Input.dispatchMouseEvent",
"params": {"type": typ, "x": x, "y": y,
"button": "left",
"clickCount": 1}}))
deadline = time.time() + timeout
while time.time() < deadline:
resp = json.loads(ws.recv())
if resp.get("id") == req_id:
break
else:
raise MouseError("mouse echo timeout for %s" % typ)
except MouseError:
raise
except Exception as e:
raise MouseError("real click failed: %s: %s"
% (type(e).__name__, e))
-30
View File
@@ -1,30 +0,0 @@
"""One module per Settings tab. Uniform: TAB, describe(ws)."""
from hatch_menu.tabs import (
connectors,
data_controls,
devices,
general,
help_support,
legal,
messaging,
permissions,
secure_store,
wallet,
)
TAB_MODULES = {
"General": general,
"Connectors": connectors,
"Wallet": wallet,
"Secure store": secure_store,
"Permissions": permissions,
"Messaging channels": messaging,
"Devices": devices,
"Data controls": data_controls,
"Help & support": help_support,
"Legal info": legal,
}
__all__ = ["TAB_MODULES", "connectors", "data_controls", "devices",
"general", "help_support", "legal", "messaging",
"permissions", "secure_store", "wallet"]
-10
View File
@@ -1,10 +0,0 @@
"""Connectors tab: read-only inventory (search + per-app Connect/View)."""
from hatch_menu import dialog
TAB = "Connectors"
def describe(ws):
"""Row inventory + text excerpt (describe-only for now)."""
return {"rows": dialog.describe_rows(ws, TAB),
"text": (dialog.dialog_text(ws) or "")[:400]}
-101
View File
@@ -1,101 +0,0 @@
"""Data controls tab: model-improvement switch (read-only otherwise).
The switch label is pinned from live recon; resolution prefers it
and falls back to single-switch, then keyword match. Sets use a
trusted click: synthetic clicks are proven no-ops here (2026-10-06).
Import/Delete rows are inventoried, never touched.
"""
import time
from approvals import cdp_evaluate
from hatch_menu import controls, dialog
from hatch_menu.mouse import real_click
TAB = "Data controls"
SWITCH_LABEL = "Help improve our AI models"
_KEYWORDS = ("improv", "train", "model", "data", "usage")
JS_AI_RECT_TMPL = """((label) => {
const d = document.querySelector('[role="dialog"]');
if (!d) return null;
const rowOf = (s) => {
let el = s.parentElement, depth = 0;
while (el && el !== d && depth < 6) {
const t = (el.innerText || '').trim();
if (t && t.length < 250) return t.toLowerCase();
el = el.parentElement;
depth += 1;
}
return '';
};
const hits = Array.from(d.querySelectorAll('[role="switch"]'))
.filter(s => rowOf(s).includes(label.toLowerCase())
|| (s.getAttribute('aria-label') || '').toLowerCase()
.includes(label.toLowerCase()));
if (hits.length !== 1) return null;
const r = hits[0].getBoundingClientRect();
return {x: r.x + r.width / 2, y: r.y + r.height / 2};
})('%s')"""
def _eval(ws, js, timeout=8.0):
try:
return cdp_evaluate(ws, js, timeout=timeout)
except Exception:
return None
def _resolve(ws):
"""The improvement switch dict, or None when not resolvable."""
if not dialog.goto_tab(ws, TAB):
return None
switches = controls.list_switches(ws)
for s in switches:
blob = ((s.get("label") or "") + " "
+ (s.get("aria") or "")).lower()
if SWITCH_LABEL.lower() in blob:
return s
if len(switches) == 1:
return switches[0]
for kw in _KEYWORDS:
for s in switches:
blob = ((s.get("label") or "") + " "
+ (s.get("aria") or "")).lower()
if kw in blob:
return s
return None
def ai_improvement(ws):
"""Improvement-switch state: True/False, None when unreadable."""
sw = _resolve(ws)
return None if sw is None else bool(sw.get("checked"))
def set_ai_improvement(ws, on):
"""Set via one trusted click; synthetic clicks are no-ops. Bool."""
sw = _resolve(ws)
if sw is None:
return False
if bool(sw.get("checked")) == bool(on):
return True
rect = _eval(ws, JS_AI_RECT_TMPL % SWITCH_LABEL)
if not rect or "x" not in rect:
return False
try:
real_click(ws, rect["x"], rect["y"])
except Exception:
return False
for _ in range(8):
time.sleep(2.0)
if ai_improvement(ws) is bool(on):
return True
return False
def describe(ws):
"""Inventory: switch state + row names (import/delete read-only)."""
state = ai_improvement(ws)
return {"ai_improvement":
("on" if state else "off") if state is not None else None,
"rows": dialog.describe_rows(ws, TAB)}
-13
View File
@@ -1,13 +0,0 @@
"""Devices tab: read-only inventory (paired devices or empty state)."""
from hatch_menu import dialog
TAB = "Devices"
def describe(ws):
"""Row inventory + empty flag (describe-only for now)."""
text = dialog.dialog_text(ws) or ""
low = text.lower()
return {"rows": dialog.describe_rows(ws, TAB),
"empty": ("no devices" in low or "don't have" in low),
"text": text[:400]}
-95
View File
@@ -1,95 +0,0 @@
"""General tab: usage balances, theme picker, redeem entrypoint.
Usage parsing moved here from invite.py (single copy). All flows are
ws-level: sessions and error shaping live in toggles.py.
"""
import re
from approvals import cdp_evaluate
from hatch_menu import controls, dialog
TAB = "General"
THEME_VALUES = ("avatar", "default", "blue", "purple", "pink",
"orange", "green", "beige", "monochrome")
_FREE_RE = re.compile(r"\bfree plan\b", re.IGNORECASE)
_PCT_RE = re.compile(r"(\d+)%\s*used")
_RESET_RE = re.compile(r"resets?\s+on\s+([A-Z][a-z]+\s+\d{1,2})",
re.IGNORECASE)
_TOK_RE = re.compile(r"\(([0-9.,]+\s*[BMK]?)\s*tokens?\s+left\)",
re.IGNORECASE)
def parse_usage_text(text):
"""Parse a General-tab usage block into a balance dict.
Returns None for empty/unreadable text. Weekly fields stay None
when the block only carries additional-tokens rows.
"""
if not text or not text.strip():
return None
out = {"plan": None, "weekly": {"pct_used": None, "resets_on": None},
"additional": {"pct_used": None, "tokens_left": None,
"never_expires": False}}
if _FREE_RE.search(text):
out["plan"] = "free"
pcts = _PCT_RE.findall(text)
if pcts:
out["weekly"]["pct_used"] = int(pcts[0])
if len(pcts) > 1:
out["additional"]["pct_used"] = int(pcts[1])
m = _RESET_RE.search(text)
if m:
out["weekly"]["resets_on"] = m.group(1)
m = _TOK_RE.search(text)
if m:
out["additional"]["tokens_left"] = m.group(1).strip()
if "never expires" in text.lower():
out["additional"]["never_expires"] = True
return out
def _eval(ws, js, timeout=8.0):
try:
return cdp_evaluate(ws, js, timeout=timeout)
except Exception:
return None
def usage(ws):
"""Usage balances from General tab. Dict, or None when unreadable."""
if not dialog.goto_tab(ws, TAB):
return None
return parse_usage_text(dialog.dialog_text(ws))
def theme(ws):
"""Current theme radio value, or None when unreadable."""
if not dialog.goto_tab(ws, TAB):
return None
for r in controls.list_radios(ws):
if r.get("checked"):
return r.get("value")
return None
def set_theme(ws, value):
"""Set theme by aria-label; verify-then-fallback. Bool."""
if not dialog.goto_tab(ws, TAB):
return False
return controls.set_radio_by_aria(ws, value)
def redeem_row(ws):
"""True when the 'Redeem invite code' row is present (eligible)."""
if not dialog.goto_tab(ws, TAB):
return False
text = (dialog.dialog_text(ws) or "").lower()
return "redeem" in text and "invite code" in text
def describe(ws):
"""Full General inventory: usage, theme, redeem entrypoint."""
return {"usage": usage(ws), "theme": theme(ws),
"redeem_row": redeem_row(ws)}
-10
View File
@@ -1,10 +0,0 @@
"""Help & support tab: read-only inventory (help links)."""
from hatch_menu import dialog
TAB = "Help & support"
def describe(ws):
"""Row inventory + text excerpt (describe-only for now)."""
return {"rows": dialog.describe_rows(ws, TAB),
"text": (dialog.dialog_text(ws) or "")[:400]}
-10
View File
@@ -1,10 +0,0 @@
"""Legal info tab: read-only inventory (legal links)."""
from hatch_menu import dialog
TAB = "Legal info"
def describe(ws):
"""Row inventory + text excerpt (describe-only for now)."""
return {"rows": dialog.describe_rows(ws, TAB),
"text": (dialog.dialog_text(ws) or "")[:400]}
-10
View File
@@ -1,10 +0,0 @@
"""Messaging channels tab: read-only inventory (channel rows)."""
from hatch_menu import dialog
TAB = "Messaging channels"
def describe(ws):
"""Row inventory + text excerpt (describe-only for now)."""
return {"rows": dialog.describe_rows(ws, TAB),
"text": (dialog.dialog_text(ws) or "")[:400]}
-422
View File
@@ -1,422 +0,0 @@
"""Permissions tab: defaults radios, website modes, protocols, advanced.
Contracts live here (single copy): toggles.py references these
constants for addressing. Drills open sub-pages, read, and come back
via the back affordance with tab re-entry fallback. All flows are
ws-level; sessions and error shaping live in toggles.py.
"""
import re
import time
from approvals import cdp_evaluate
from hatch_menu import controls, dialog
from hatch_menu.mouse import escape, real_click
TAB = "Permissions"
ROOT_MARK = "Manage permissions"
CONNECTOR_HEADING = "Connector defaults"
WEB_HEADING = "Web access defaults"
DEFAULT_VALUES = ("auto_allow", "always_ask")
WEBSITE_MODES = ("Allow", "Ask", "Deny")
ADV_LABELS = {"transparent_proxy": "Transparent proxy",
"tls_interception": "TLS interception",
"sni_mismatch_rejection": "SNI mismatch rejection"}
# Row titles (first line of each protocol row) pinned live 2026-10-06:
# network primitives on every node checked; MCP titles kept
# defensively in case they appear on other plans/accounts.
PROTOCOL_SLUGS = {"Outbound SSH": "outbound-ssh",
"Outgoing email (SMTP)": "smtp",
"Email mailbox access (IMAP, POP3)": "imap-pop3",
"Database connections": "database",
"File transfer (FTP)": "ftp",
"External DNS lookups": "dns",
"Other TCP connections": "other-tcp",
"Other UDP traffic": "other-udp",
"Model Context Protocol servers (SSE)": "mcp-sse",
"Model Context Protocol servers (Streamable HTTP)":
"mcp-streamable",
"Agent Skills endpoints": "agent-skills",
"MCP Apps (UI extensions)": "mcp-apps",
"MCP remote OAuth": "mcp-oauth"}
JS_WEBSITES = """(() => {
const d = document.querySelector('[role="dialog"]');
if (!d) return null;
const out = [];
for (const b of d.querySelectorAll('button')) {
const m = (b.getAttribute('aria-label') || '').match(
/^Change permission mode for (.+),\\s*(Allow|Ask|Deny)$/i);
if (!m) continue;
const r = b.getBoundingClientRect();
out.push({host: m[1].trim(), mode: m[2],
x: r.x + r.width / 2, y: r.y + r.height / 2});
}
return out;
})()"""
JS_MODE_MENU = """(() => {
return Array.from(document.querySelectorAll('[role="menuitem"]'))
.map(m => ({text: (m.innerText || '').trim()}));
})()"""
JS_CLICK_MODE = """((mode) => {
const m = Array.from(document.querySelectorAll('[role="menuitem"]'))
.find(el => (el.innerText || '').trim() === mode);
if (!m) return 'NO_MATCH';
m.click();
return 'CLICKED';
})('%s')"""
JS_PROTO_ROWS = """(() => {
const d = document.querySelector('[role="dialog"]');
if (!d) return null;
return Array.from(d.querySelectorAll('[role="switch"]')).map(s => {
let el = s.parentElement, title = '', depth = 0;
while (el && el !== d && depth < 6) {
const t = (el.innerText || '').trim().split('\\n')[0] || '';
if (t) { title = t.slice(0, 80); break; }
el = el.parentElement;
depth += 1;
}
const r = s.getBoundingClientRect();
return {title: title,
checked: s.getAttribute('aria-checked') === 'true',
x: r.x + r.width / 2, y: r.y + r.height / 2};
});
})()"""
def _eval(ws, js, timeout=8.0):
try:
return cdp_evaluate(ws, js, timeout=timeout)
except Exception:
return None
def _stable_rows(ws, js, retries=4, pause=1.5):
"""Repeat a row read until two consecutive reads agree.
Guards mid-animation partial DOM (innerText shifts while the
sub-page slides in). Returns the agreed list, or None.
"""
last = "sentinel"
for _ in range(retries):
rows = _eval(ws, js)
if isinstance(rows, list) and rows == last:
return rows
last = rows if isinstance(rows, list) else "sentinel"
time.sleep(pause)
return last if isinstance(last, list) else None
def _slug(title):
"""Protocol slug: registry hit, else slugified, else None."""
if title in PROTOCOL_SLUGS:
return PROTOCOL_SLUGS[title]
clean = re.sub(r"[^a-z0-9]+", "-",
title.strip().lower()).strip("-")
return clean or None
def _canon_mode(mode):
"""Canonical Allow/Ask/Deny (case-insensitive); passthrough else."""
for m in WEBSITE_MODES:
if (mode or "").lower() == m.lower():
return m
return mode
def resolve_protocol(name):
"""Slug/title -> row title, None when unresolvable.
Exact slug or title first; then a unique case-insensitive
substring over titles+slugs (so 'ssh' finds Outbound SSH).
"""
if not isinstance(name, str) or not name.strip():
return None
want = name.strip().lower()
for title, slug in PROTOCOL_SLUGS.items():
if want == slug or want == title.lower():
return title
hits = [t for t, s in PROTOCOL_SLUGS.items()
if want in t.lower() or want in s]
if len(hits) == 1:
return hits[0]
return None
def _back_to_root(ws):
"""Back affordance, else tab re-entry; verify root text."""
dialog.go_back(ws)
if ROOT_MARK in (dialog.dialog_text(ws) or ""):
return True
if not dialog.goto_tab(ws, TAB):
return False
return ROOT_MARK in (dialog.dialog_text(ws) or "")
def defaults(ws):
"""Connector + web default values (each value or None)."""
if not dialog.goto_tab(ws, TAB):
return {"connector_defaults": None, "web_access": None}
heads = {CONNECTOR_HEADING.lower(): "connector_defaults",
WEB_HEADING.lower(): "web_access"}
vals = {CONNECTOR_HEADING.lower(): [], WEB_HEADING.lower(): []}
for r in controls.list_radios(ws):
h = (r.get("heading") or "").lower()
if h in vals and r.get("checked"):
vals[h].append(r.get("value"))
out = {}
for h, key in heads.items():
out[key] = vals[h][0] if len(vals[h]) == 1 else None
return out
def set_default(ws, which, value):
"""Set one defaults radio. Bool."""
if not dialog.goto_tab(ws, TAB):
return False
heading = CONNECTOR_HEADING if which == "connector_defaults" \
else WEB_HEADING
return controls.set_radio_by_heading(ws, heading, value)
def ensure_advanced(ws):
"""Expand Advanced network settings when collapsed. Bool."""
if not dialog.goto_tab(ws, TAB):
return False
labels = [s.get("label", "") for s in controls.list_switches(ws)]
if any("Transparent proxy" in lab for lab in labels):
return True
if not dialog.click_row(ws, "Advanced network settings", TAB):
return False
time.sleep(0.6)
labels = [s.get("label", "") for s in controls.list_switches(ws)]
return any("Transparent proxy" in lab for lab in labels)
def advanced(ws):
"""Advanced switch states {key: on/off/None}."""
if not ensure_advanced(ws):
return {k: None for k in ADV_LABELS}
out = {}
for key, label in ADV_LABELS.items():
state = None
for s in controls.list_switches(ws):
if label.lower() in (s.get("label") or "").lower():
state = bool(s.get("checked"))
break
out[key] = ("on" if state else "off") if state is not None \
else None
return out
def set_advanced(ws, key, on):
"""Set one advanced switch. Bool."""
if key not in ADV_LABELS or not ensure_advanced(ws):
return False
return controls.set_switch(ws, ADV_LABELS[key], on)
def _websites_raw(ws):
"""Drill into Websites; rows or None (stays on sub-page)."""
if not dialog.click_row(ws, "Websites", TAB):
return None
return _stable_rows(ws, JS_WEBSITES)
def websites(ws):
"""[{host, mode}] (back at root afterwards)."""
rows = _websites_raw(ws)
if rows is None:
return []
out = [{"host": r.get("host"), "mode": _canon_mode(r.get("mode"))}
for r in rows]
_back_to_root(ws)
return out
def website_mode(ws, host):
"""Mode for one host, or None when absent/unreadable."""
for row in websites(ws):
if (row.get("host") or "").lower() == host.lower():
return row.get("mode")
return None
def set_website_mode(ws, host, mode):
"""Set one host mode via the mode chooser. Bool.
One-way for Ask/Deny: the override row leaves the allowed list
(no add UI), so removal verifies by absence. No-op when already
there; absent hosts fail (nothing to click).
"""
if mode not in WEBSITE_MODES:
return False
rows = _websites_raw(ws)
if rows is None:
return False
target = next((r for r in rows
if (r.get("host") or "").lower() == host.lower()),
None)
if target is None:
_back_to_root(ws)
return False
if _canon_mode(target.get("mode")) == mode:
_back_to_root(ws)
return True
try:
real_click(ws, target["x"], target["y"])
except Exception:
_back_to_root(ws)
return False
time.sleep(0.8)
items = _eval(ws, JS_MODE_MENU)
texts = [(i.get("text") or "") for i in items] \
if isinstance(items, list) else []
if mode not in texts:
escape(ws)
_back_to_root(ws)
return False
if _eval(ws, JS_CLICK_MODE % mode) != "CLICKED":
escape(ws)
_back_to_root(ws)
return False
for _ in range(5):
time.sleep(2.0)
rows = _eval(ws, JS_WEBSITES)
if not isinstance(rows, list):
continue
cur = next((_canon_mode(r.get("mode")) for r in rows
if (r.get("host") or "").lower() == host.lower()),
None)
if mode in ("Ask", "Deny"):
if cur is None:
_back_to_root(ws)
return True
elif cur == mode:
_back_to_root(ws)
return True
escape(ws)
_back_to_root(ws)
return False
def _protocols_raw(ws):
"""Drill into protocols; rows or None (stays on sub-page)."""
if not dialog.click_row(ws, "Direct network protocols", TAB):
return None
return _stable_rows(ws, JS_PROTO_ROWS)
def protocols(ws):
"""[{slug, title, on}] (back at root afterwards)."""
rows = _protocols_raw(ws)
if rows is None:
return []
out = [{"slug": _slug(r.get("title", "")),
"title": r.get("title", ""),
"on": "on" if r.get("checked") else "off"} for r in rows]
_back_to_root(ws)
return out
def protocol_state(ws, title):
"""on/off for one protocol row title, None when absent."""
for row in protocols(ws):
if row.get("title") == title:
return row.get("on")
return None
def set_protocol(ws, title, on):
"""Set one protocol switch in place; readback before returning."""
rows = _protocols_raw(ws)
if rows is None:
return False
target = next((r for r in rows if r.get("title") == title), None)
if target is None:
_back_to_root(ws)
return False
want = bool(on)
if bool(target.get("checked")) == want:
_back_to_root(ws)
return True
try:
real_click(ws, target["x"], target["y"])
except Exception:
_back_to_root(ws)
return False
for _ in range(8):
time.sleep(2.0)
rows = _eval(ws, JS_PROTO_ROWS)
if not isinstance(rows, list):
continue
cur = next((r for r in rows if r.get("title") == title), None)
if cur is not None and bool(cur.get("checked")) == want:
_back_to_root(ws)
return True
_back_to_root(ws)
# In-dialog verify missed (slow commit or commit-on-close); the
# toggles-level fresh readback is the source of truth.
return False
def manage_counts(ws):
"""Manageable-row summary counts (name -> count|None)."""
if not dialog.goto_tab(ws, TAB):
return {}
text = dialog.dialog_text(ws) or ""
out = {}
for name in ("Websites", "Connectors", "Scheduled tasks",
"Direct network protocols"):
m = re.search(re.escape(name) + r"\s*(\d+)", text)
out[name] = int(m.group(1)) if m else None
return out
def scheduled_tasks(ws):
"""[{name, cadence}] (back at root afterwards; empty when none)."""
if not dialog.click_row(ws, "Scheduled tasks", TAB):
return []
time.sleep(0.6)
text = dialog.dialog_text(ws) or ""
_back_to_root(ws)
rows = []
lines = [line.strip() for line in text.splitlines()
if line.strip()]
for i, line in enumerate(lines):
if re.search(r"\b(daily|weekly|hourly|every|min)\b", line,
re.IGNORECASE) and i > 0:
rows.append({"name": lines[i - 1], "cadence": line})
return rows
def all_toggles(ws):
"""Flat map of every settable Permissions toggle (for list)."""
out = {}
defs = defaults(ws)
out["permissions.connector_defaults"] = defs.get("connector_defaults")
out["permissions.web_access"] = defs.get("web_access")
adv = advanced(ws)
for key, val in adv.items():
out["permissions.advanced." + key] = val
for row in protocols(ws):
if row.get("slug"):
out["permissions.protocols:" + row["slug"]] = row.get("on")
for row in websites(ws):
if row.get("host"):
out["permissions.websites:" + row["host"]] = row.get("mode")
return out
def describe(ws):
"""Full Permissions inventory: defaults, counts, adv, rows."""
return {"defaults": defaults(ws), "counts": manage_counts(ws),
"advanced": advanced(ws), "websites": websites(ws),
"protocols": protocols(ws),
"scheduled_tasks": scheduled_tasks(ws)}
-10
View File
@@ -1,10 +0,0 @@
"""Secure store tab: read-only inventory (secret entries + add)."""
from hatch_menu import dialog
TAB = "Secure store"
def describe(ws):
"""Row inventory + text excerpt (describe-only for now)."""
return {"rows": dialog.describe_rows(ws, TAB),
"text": (dialog.dialog_text(ws) or "")[:400]}
-10
View File
@@ -1,10 +0,0 @@
"""Wallet tab: read-only inventory (payment methods + add row)."""
from hatch_menu import dialog
TAB = "Wallet"
def describe(ws):
"""Row inventory + text excerpt (describe-only for now)."""
return {"rows": dialog.describe_rows(ws, TAB),
"text": (dialog.dialog_text(ws) or "")[:400]}
-319
View File
@@ -1,319 +0,0 @@
"""Toggle registry + sessions over controls.py and tabs/*.
Every settable toggle has ONE address; contracts live in the tab
modules (single-copy per part), addressing here. Set flows read back
through a fresh session and never partially report success.
Caller errors (unknown node/toggle/value) raise MenuError before any
CDP traffic. Transport failures return {"ok": False, ...}.
"""
from approvals import VALID_NODES, get_cdp_ws
from hatch_menu import controls, dialog
from hatch_menu.mouse import close, escape
from hatch_menu.tabs import TAB_MODULES
_PERM = TAB_MODULES["Permissions"]
_DC = TAB_MODULES["Data controls"]
_GEN = TAB_MODULES["General"]
class MenuError(ValueError):
"""Caller error: unknown node, toggle, or value. Raised before CDP."""
TOGGLES = {
"permissions.connector_defaults": {
"tab": "Permissions", "kind": "radio-heading",
"heading": _PERM.CONNECTOR_HEADING,
"values": _PERM.DEFAULT_VALUES},
"permissions.web_access": {
"tab": "Permissions", "kind": "radio-heading",
"heading": _PERM.WEB_HEADING,
"values": _PERM.DEFAULT_VALUES},
"permissions.advanced.transparent_proxy": {
"tab": "Permissions", "kind": "adv-switch",
"label": _PERM.ADV_LABELS["transparent_proxy"],
"values": ("on", "off")},
"permissions.advanced.tls_interception": {
"tab": "Permissions", "kind": "adv-switch",
"label": _PERM.ADV_LABELS["tls_interception"],
"values": ("on", "off")},
"permissions.advanced.sni_mismatch_rejection": {
"tab": "Permissions", "kind": "adv-switch",
"label": _PERM.ADV_LABELS["sni_mismatch_rejection"],
"values": ("on", "off")},
"data_controls.ai_improvement": {
"tab": "Data controls", "kind": "switch",
"values": ("on", "off"),
# Live 2026-10-06: the site ignores every input gesture here
# (synthetic/real/double/hover/keyboard/drag) — reads fine.
"readonly": True},
"general.theme": {
"tab": "General", "kind": "radio-aria",
"values": _GEN.THEME_VALUES},
}
WEBSITE_PREFIX = "permissions.websites:"
PROTOCOL_PREFIX = "permissions.protocols:"
def _check_node(node):
if node not in VALID_NODES:
raise MenuError("unknown node: %r (valid: %s)"
% (node, ", ".join(VALID_NODES)))
def normalize_onoff(value):
"""on/off/true/false/1/0/yes/no -> bool. None when invalid."""
if isinstance(value, bool):
return value
if not isinstance(value, str):
return None
v = value.strip().lower()
if v in ("on", "true", "1", "yes"):
return True
if v in ("off", "false", "0", "no"):
return False
return None
def resolve_toggle(name):
"""Resolve a toggle address to its spec. Raises MenuError."""
if not isinstance(name, str) or not name:
raise MenuError("toggle name must be a non-empty string")
if name in TOGGLES:
spec = dict(TOGGLES[name])
spec["name"] = name
return spec
if name.startswith(WEBSITE_PREFIX) and len(name) > len(WEBSITE_PREFIX):
return {"name": name, "tab": "Permissions", "kind": "website",
"host": name[len(WEBSITE_PREFIX):],
"values": _PERM.WEBSITE_MODES}
if name.startswith(PROTOCOL_PREFIX) and len(name) > len(PROTOCOL_PREFIX):
label = _PERM.resolve_protocol(name[len(PROTOCOL_PREFIX):])
if label is None:
raise MenuError("unknown protocol: %r (see list_toggles)"
% (name,))
return {"name": name, "tab": "Permissions", "kind": "protocol",
"label": label, "values": ("on", "off")}
raise MenuError("unknown toggle: %r (see list_toggles)" % (name,))
def _session(node):
_check_node(node)
try:
ws, _ = get_cdp_ws(node)
except Exception as e:
raise MenuError("CDP unreachable for %s: %s" % (node, e))
return ws
def _normalize_value(spec, value):
kind = spec["kind"]
if kind == "website":
if not isinstance(value, str):
raise MenuError("mode must be one of %s"
% (spec["values"],))
for m in spec["values"]:
if value.strip().lower() == m.lower():
return m
raise MenuError("mode must be one of %s (got %r)"
% (spec["values"], value))
if kind in ("switch", "adv-switch", "protocol"):
b = normalize_onoff(value)
if b is None:
raise MenuError("value must be on/off (got %r)" % (value,))
return "on" if b else "off"
if kind == "radio-heading":
if value not in spec["values"]:
raise MenuError("%s must be one of %s (got %r)"
% (spec["name"], spec["values"], value))
return value
if kind == "radio-aria":
if not isinstance(value, str):
raise MenuError("theme must be one of %s" % (spec["values"],))
v = value.strip().lower()
if v in spec["values"]:
return v
raise MenuError("theme must be one of %s (got %r)"
% (spec["values"], value))
raise MenuError("cannot set kind %r" % (kind,))
def get_toggle(node, name):
"""Read one toggle. Returns {"ok", "node", "toggle", "value"}."""
spec = resolve_toggle(name)
_check_node(node)
try:
ws = _session(node)
except MenuError as e:
return {"ok": False, "node": node, "toggle": name,
"error": str(e)}
try:
if not dialog.open_settings(ws):
return {"ok": False, "node": node, "toggle": name,
"error": "settings dialog did not open"}
if not dialog.goto_tab(ws, spec["tab"]):
return {"ok": False, "node": node, "toggle": name,
"error": "tab did not open: %s" % spec["tab"]}
kind = spec["kind"]
if kind == "radio-heading":
vals = {r["value"]: r["checked"]
for r in controls.list_radios(ws)
if (r.get("heading") or "").lower()
== spec["heading"].lower()}
on = [v for v, c in vals.items() if c]
value = on[0] if len(on) == 1 else None
elif kind == "radio-aria":
vals = [(r.get("value"), r.get("checked"))
for r in controls.list_radios(ws)]
on = [v for v, c in vals if c]
value = on[0] if len(on) == 1 else None
elif kind == "switch":
state = _DC.ai_improvement(ws)
value = ("on" if state else "off") if state is not None \
else None
elif kind == "adv-switch":
_PERM.ensure_advanced(ws)
state = controls.switch_state(ws, spec["label"])
value = ("on" if state else "off") if state is not None \
else None
elif kind == "website":
value = _PERM.website_mode(ws, spec["host"])
elif kind == "protocol":
value = _PERM.protocol_state(ws, spec["label"])
else:
value = None
if value is None:
if kind == "website":
return {"ok": False, "node": node, "toggle": name,
"error": "host not in Websites list (effective: "
"permissions.web_access default)"}
return {"ok": False, "node": node, "toggle": name,
"error": "toggle not readable (site changed?)"}
return {"ok": True, "node": node, "toggle": name, "value": value}
except Exception as e:
return {"ok": False, "node": node, "toggle": name,
"error": "%s: %s" % (type(e).__name__, e)}
finally:
escape(ws)
close(ws)
def set_toggle(node, name, value):
"""Set one toggle with readback. Never partially reports success."""
spec = resolve_toggle(name)
if spec.get("readonly"):
raise MenuError("%s is read-only: the site ignores all input "
"gestures there (locked?)" % name)
want = _normalize_value(spec, value)
_check_node(node)
try:
ws = _session(node)
except MenuError as e:
return {"ok": False, "node": node, "toggle": name,
"error": str(e)}
try:
if not dialog.open_settings(ws):
return {"ok": False, "node": node, "toggle": name,
"error": "settings dialog did not open"}
if not dialog.goto_tab(ws, spec["tab"]):
return {"ok": False, "node": node, "toggle": name,
"error": "tab did not open: %s" % spec["tab"]}
kind = spec["kind"]
if kind == "radio-heading":
ok = controls.set_radio_by_heading(ws, spec["heading"], want)
elif kind == "radio-aria":
ok = controls.set_radio_by_aria(ws, want)
elif kind == "switch":
ok = _DC.set_ai_improvement(ws, want == "on")
elif kind == "adv-switch":
_PERM.ensure_advanced(ws)
ok = controls.set_switch(ws, spec["label"], want == "on")
elif kind == "website":
ok = _PERM.set_website_mode(ws, spec["host"], want)
elif kind == "protocol":
ok = _PERM.set_protocol(ws, spec["label"], want == "on")
else:
ok = False
# The fresh-session readback is the source of truth: switch
# commits can land slowly or on dialog close, after the
# in-flow verify had its chance.
readback = get_toggle(node, name)
if readback.get("ok") and readback.get("value") == want:
out = {"ok": True, "node": node, "toggle": name,
"value": want}
if not ok:
out["readback_only"] = True
return out
return {"ok": False, "node": node, "toggle": name,
"error": "readback mismatch (want %r, got %r)"
% (want, readback.get("value"))}
except Exception as e:
return {"ok": False, "node": node, "toggle": name,
"error": "%s: %s" % (type(e).__name__, e)}
finally:
escape(ws)
close(ws)
def list_toggles(node, tab=None):
"""All toggle states, optionally filtered to one tab."""
_check_node(node)
if tab is not None and tab not in TAB_MODULES:
raise MenuError("unknown tab: %r (valid: %s)"
% (tab, sorted(TAB_MODULES)))
want = [tab] if tab else ["Permissions", "Data controls", "General"]
try:
ws = _session(node)
except MenuError as e:
return {"ok": False, "node": node, "error": str(e)}
try:
if not dialog.open_settings(ws):
return {"ok": False, "node": node,
"error": "settings dialog did not open"}
out = {}
if "Permissions" in want:
dialog.goto_tab(ws, "Permissions")
out.update(_PERM.all_toggles(ws))
if "Data controls" in want:
state = _DC.ai_improvement(ws)
out["data_controls.ai_improvement"] = \
("on" if state else "off") if state is not None else None
if "General" in want:
dialog.goto_tab(ws, "General")
out["general.theme"] = _GEN.theme(ws)
return {"ok": True, "node": node, "toggles": out}
except Exception as e:
return {"ok": False, "node": node,
"error": "%s: %s" % (type(e).__name__, e)}
finally:
escape(ws)
close(ws)
def describe_tab(node, tab):
"""Full inventory of one tab (debugging/patching aid)."""
_check_node(node)
if tab not in TAB_MODULES:
raise MenuError("unknown tab: %r (valid: %s)"
% (tab, sorted(TAB_MODULES)))
try:
ws = _session(node)
except MenuError as e:
return {"ok": False, "node": node, "tab": tab, "error": str(e)}
try:
if not dialog.open_settings(ws):
return {"ok": False, "node": node, "tab": tab,
"error": "settings dialog did not open"}
if not dialog.goto_tab(ws, tab):
return {"ok": False, "node": node, "tab": tab,
"error": "tab did not open"}
return {"ok": True, "node": node, "tab": tab,
"inventory": TAB_MODULES[tab].describe(ws)}
except Exception as e:
return {"ok": False, "node": node, "tab": tab,
"error": "%s: %s" % (type(e).__name__, e)}
finally:
escape(ws)
close(ws)
-348
View File
@@ -1,348 +0,0 @@
#!/usr/bin/env python3
"""Host-side fleet evidence for network/PID-blind shells.
`box fleet status` probes each node live (pgrep for the chromium process,
HTTP to the CDP relay on the peer IP). Both probes assume the caller's
network + PID namespace is the bl host's. From a sandboxed shell (own PID
and net namespaces, no sudo, no route to 10.201.x.x) both probes always
fail, so every node misreports as STOPPED even with a healthy fleet.
This module provides the fallback signal: evidence written by the
host-side watchdogs that run on bl unsandboxed via systemd timers:
- cdp-relay-watchdog (every 5 min, all active registry nodes):
log ``cdp-relay-watchdog.log`` + journal unit
``cdp-relay-watchdog.service``. Proves the CDP relay path end to end.
- chromebox-watchdog (every 2 min, same nodes, one timer per node):
log ``chromebox-watchdog.log`` + journal units
``chromebox-watchdog@<node>.service``. Curls CDP inside the node netns,
so it proves browser + in-netns CDP.
- ``chromebox-<node>.log`` mtime: chromium's own stdout. Fresh output
proves the browser process is alive. Used only for nodes outside
watchdog coverage — and only to conclude "alive", never "dead".
Both watchdogs are silent-when-healthy: a failure is ALWAYS logged, so a
recent timer run (journal "Starting" line) with no newer failure line for
the node means that run found the node healthy.
Verdicts: "healthy" | "degraded" | "down" | "unknown".
"""
import json
import os
import re
import subprocess
import time
from datetime import datetime, timezone
from pathlib import Path
NETVM_ROOT = Path("/home/super/Projects/NetVM")
RELAY_LOG = NETVM_ROOT / "cdp-relay-watchdog.log"
CHROMEBOX_LOG = NETVM_ROOT / "chromebox-watchdog.log"
ALL_NODES = ("muse", "pip", "646", "opm", "def", "dev")
def _covered_nodes():
"""Nodes with watchdog coverage, from the fleet registry.
Both watchdogs supervise every active registry node. Falls back to
ALL_NODES when the registry is unreadable, so a broken registry can
never silently narrow fleet status to a subset of the fleet.
"""
try:
import importlib.util
spec = importlib.util.spec_from_file_location(
"netvm_registry", NETVM_ROOT / "bin" / "netvm-registry.py")
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return tuple(sorted(mod.active_nodes()))
except Exception:
return ALL_NODES
RELAY_NODES = _covered_nodes()
CHROMEBOX_NODES = _covered_nodes()
RELAY_UNIT = "cdp-relay-watchdog.service"
CHROMEBOX_UNIT_TMPL = "chromebox-watchdog@{node}.service"
# A watchdog run older than this proves nothing (timer may be dead).
RELAY_STALE_MIN = 15
CHROMEBOX_STALE_MIN = 8
# Chromium stdout older than this proves nothing (idle browsers go quiet).
CHROME_LOG_FRESH_MIN = 20
_LOG_TS_RE = re.compile(r"^\[(\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2})Z\]")
def _utcnow():
return datetime.now(timezone.utc)
def parse_log_ts(line):
"""Parse a ``[YYYY-MM-DDTHH:MM:SSZ]`` log prefix. None if absent."""
m = _LOG_TS_RE.match(line)
if not m:
return None
try:
return datetime.strptime(m.group(1), "%Y-%m-%dT%H:%M:%S").replace(
tzinfo=timezone.utc)
except ValueError:
return None
def classify_chromebox_line(line):
"""Classify one chromebox-watchdog log line.
Returns "healthy" | "degraded" | "down", or None when the line
carries no verdict (rotation markers, relay stdout passthrough).
"""
if "log rotated" in line:
return None
if "relaunch FAILED" in line:
return "down"
if "relaunch OK" in line:
return "healthy"
if "recovered" in line and "relaunch not needed" in line:
return "healthy"
if "not healthy yet" in line:
return "degraded"
if "proceeding with chrome relaunch" in line:
return "degraded"
if "relaunching chromebox" in line:
return "degraded"
if "warp partition detected" in line:
return "degraded"
if "skipping relaunch (probably still starting)" in line:
return "degraded"
return None
def classify_relay_line(line):
"""Classify one cdp-relay-watchdog log line (None = no verdict)."""
if "relay restart FAILED" in line:
return "down"
if "FAIL_LOUD" in line:
return "down"
if "relay restarted OK" in line:
return "healthy"
if "relay unhealthy" in line and "restarting" in line:
# Always followed by an OK/FAILED line; a trailing one means the
# restart crashed mid-flight.
return "down"
return None
def last_verdict(lines, node, classify):
"""Newest (verdict, ts, line) for ``[node]``. None if no verdict line."""
tag = "[%s]" % node
best = None
for line in lines:
if tag not in line:
continue
verdict = classify(line)
if verdict is None:
continue
ts = parse_log_ts(line)
if ts is None:
continue
if best is None or ts >= best[1]:
best = (verdict, ts, line.strip()[:160])
return best
def _tail_lines(path, max_bytes=65536):
try:
size = os.path.getsize(path)
with open(path, "rb") as f:
if size > max_bytes:
f.seek(size - max_bytes)
f.readline() # drop partial first line
return f.read().decode("utf-8", errors="replace").splitlines()
except OSError:
return []
def query_journal_starts(units, since_min=25, timeout=20):
"""Map each unit -> newest run-start (aware UTC). Missing on failure.
Uses ``-o json``: the short-format "Starting" line carries the unit
description, not the unit name, so exact per-unit matching needs the
structured UNIT field.
"""
cmd = ["journalctl", "--no-pager", "-o", "json",
"--since", "%d min ago" % since_min]
for u in units:
cmd.extend(["-u", u])
try:
r = subprocess.run(cmd, capture_output=True, text=True, timeout=timeout)
except (OSError, subprocess.TimeoutExpired):
return {}
if r.returncode != 0:
return {}
want = set(units)
starts = {}
for line in (r.stdout or "").splitlines():
try:
e = json.loads(line)
except ValueError:
continue
if e.get("UNIT") not in want:
continue
if not (e.get("MESSAGE") or "").startswith("Starting"):
continue
try:
ts = datetime.fromtimestamp(
int(e["__REALTIME_TIMESTAMP"]) / 1e6, tz=timezone.utc)
except (KeyError, ValueError, TypeError, OverflowError):
continue
u = e["UNIT"]
if u not in starts or ts > starts[u]:
starts[u] = ts
return starts
def _verdict_since_run(verdict_row, run_ts):
"""True when the verdict line is newer than (or from) the last run."""
if verdict_row is None or run_ts is None:
return False
return verdict_row[1] >= run_ts
def browser_verdict(node, chromebox_lines, run_ts, chrome_log_mtime=None,
now=None):
"""(verdict, detail) for the node's browser process."""
now = now or _utcnow()
row = last_verdict(chromebox_lines, node, classify_chromebox_line)
if node in CHROMEBOX_NODES:
if run_ts is None:
return ("unknown", "no chromebox-watchdog run in journal window")
if (now - run_ts).total_seconds() > CHROMEBOX_STALE_MIN * 60:
return ("unknown", "chromebox-watchdog run is stale")
if _verdict_since_run(row, run_ts):
return (row[0], "watchdog: %s" % row[2])
return ("healthy", "watchdog run silent (silent-when-healthy)")
# Nodes outside watchdog coverage: chromium stdout proves alive only.
if chrome_log_mtime is not None and (
now - chrome_log_mtime).total_seconds() < CHROME_LOG_FRESH_MIN * 60:
return ("healthy", "chromebox-%s.log fresh" % node)
return ("unknown", "no watchdog coverage for %s" % node)
def cdp_verdict(node, relay_lines, run_ts, now=None):
"""(verdict, detail) for the node's host-reachable CDP relay path."""
now = now or _utcnow()
if node not in RELAY_NODES:
return ("unknown", "no relay-monitor coverage for %s" % node)
if run_ts is None:
return ("unknown", "no cdp-relay-watchdog run in journal window")
if (now - run_ts).total_seconds() > RELAY_STALE_MIN * 60:
return ("unknown", "cdp-relay-watchdog run is stale")
row = last_verdict(relay_lines, node, classify_relay_line)
if _verdict_since_run(row, run_ts):
return (row[0], "relay watchdog: %s" % row[2])
return ("healthy", "relay watchdog run silent (silent-when-healthy)")
def chrome_log_mtime(node):
"""Mtime of chromium's stdout log as aware UTC. None if missing."""
try:
return datetime.fromtimestamp(
os.path.getmtime(NETVM_ROOT / ("chromebox-%s.log" % node)),
tz=timezone.utc)
except OSError:
return None
_CACHE = {"at": 0.0, "nodes": frozenset(), "data": {}}
_CACHE_TTL_S = 60
def collect(nodes=None, _journal_starts=None, _relay_lines=None,
_chromebox_lines=None, _chrome_mtimes=None, _now=None):
"""Per-node host evidence. Underscore args are seams for tests."""
nodes = list(nodes or ALL_NODES)
live = (_journal_starts is None and _relay_lines is None
and _chromebox_lines is None and _chrome_mtimes is None
and _now is None)
if live:
key = frozenset(nodes)
if (key <= _CACHE["nodes"]
and time.monotonic() - _CACHE["at"] < _CACHE_TTL_S):
return {n: _CACHE["data"][n] for n in nodes if n in _CACHE["data"]}
now = _now or _utcnow()
if _journal_starts is None:
units = [RELAY_UNIT] + [CHROMEBOX_UNIT_TMPL.format(node=n)
for n in nodes if n in CHROMEBOX_NODES]
_journal_starts = query_journal_starts(units)
if _relay_lines is None:
_relay_lines = _tail_lines(RELAY_LOG)
if _chromebox_lines is None:
_chromebox_lines = _tail_lines(CHROMEBOX_LOG)
out = {}
for node in nodes:
if _chrome_mtimes is not None and node in _chrome_mtimes:
mtime = _chrome_mtimes[node]
else:
mtime = chrome_log_mtime(node) if node not in CHROMEBOX_NODES else None
b_verd, b_det = browser_verdict(
node, _chromebox_lines,
_journal_starts.get(CHROMEBOX_UNIT_TMPL.format(node=node)),
chrome_log_mtime=mtime, now=now)
c_verd, c_det = cdp_verdict(
node, _relay_lines, _journal_starts.get(RELAY_UNIT), now=now)
out[node] = {
"browser": b_verd,
"browser_detail": b_det,
"cdp": c_verd,
"cdp_detail": c_det,
}
if live:
_CACHE["at"] = time.monotonic()
_CACHE["nodes"] = frozenset(nodes)
_CACHE["data"] = out
return out
def effective_status(local_proc, local_cdp, browser_v, cdp_v):
"""Map (local probes, host verdicts) -> (status, source).
Host evidence only ever overrides the fully-blind pattern (both
local probes negative — the sandbox signature). It never overrides
a live local signal, so a fresh outage on the host always wins.
"""
if local_cdp:
# CDP answers: the browser is definitionally alive.
return ("ACTIVE", "local")
if local_proc:
return ("CDP_DOWN", "local")
# Both local probes negative: consult host evidence.
if browser_v == "down":
return ("STOPPED", "host-evidence")
if browser_v == "unknown":
return ("UNKNOWN", "host-evidence")
if cdp_v == "healthy":
return ("ACTIVE", "host-evidence")
if cdp_v == "down":
return ("CDP_DOWN", "host-evidence")
return ("UNKNOWN", "host-evidence")
def main(argv=None):
import argparse
ap = argparse.ArgumentParser(description="Show host-side fleet evidence")
ap.add_argument("--json", action="store_true")
args = ap.parse_args(argv)
data = collect()
if args.json:
print(json.dumps({"ok": True, "evidence": data}, indent=2))
return
for node, ev in data.items():
print("%-6s browser=%-8s cdp=%-8s" % (node, ev["browser"], ev["cdp"]))
print(" browser: %s" % ev["browser_detail"])
print(" cdp: %s" % ev["cdp_detail"])
if __name__ == "__main__":
main()
-70
View File
@@ -1,70 +0,0 @@
#!/bin/bash
# identity-audit-check.sh - Check VM identity audit for drift
#
# Lives on bl (/home/super/Projects/NetVM/bin/). Reads a bl-local cached
# copy of the VM's identity audit JSON and reports drift.
#
# Why a cache: bl cannot SSH to the VM (VM only accepts the operator
# container's key). The VM's hourly audit cron (/srv/board/bin/identity-audit.py)
# should scp /srv/board/data/identity-audit.json to bl at:
# /home/super/Projects/NetVM/var/identity-audit.json
# after each run. Until that push is wired, the cache is refreshed manually
# or by the identity-audit-watch cron.
#
# Called by:
# - the identity-audit-watch cron (replaces inline SSH one-liner)
# - box-ctl.py `identity-audit` action
#
# Exit codes:
# 0 - clean (no drift; warnings are expected and silent)
# 1 - drift detected (items printed to stdout, one per line)
# 2 - audit cache missing/unreadable (needs attention)
#
# Output on drift: one drift item per line, prefixed with "DRIFT: "
set -u
CACHE_PATH="/home/super/Projects/NetVM/var/identity-audit.json"
# Allow override for testing
if [ $# -ge 1 ] && [ -f "$1" ]; then
CACHE_PATH="$1"
fi
if [ ! -f "$CACHE_PATH" ]; then
echo "ERROR: identity audit cache missing: $CACHE_PATH" >&2
echo "The VM hourly audit should push /srv/board/data/identity-audit.json here after each run." >&2
exit 2
fi
audit_json=$(cat "$CACHE_PATH" 2>/dev/null)
if [ -z "$audit_json" ]; then
echo "ERROR: identity audit cache unreadable: $CACHE_PATH" >&2
exit 2
fi
# Parse with python3
drift_output=$(printf '%s' "$audit_json" | python3 -c "
import json, sys
try:
data = json.load(sys.stdin)
except Exception as e:
print('ERROR: invalid JSON in audit cache: %s' % e, file=sys.stderr)
sys.exit(2)
for item in data.get('drift', []):
print('DRIFT: ' + str(item))
" 2>&1)
parse_rc=$?
if [ $parse_rc -eq 2 ]; then
echo "$drift_output" >&2
exit 2
fi
if [ -n "$drift_output" ]; then
echo "$drift_output"
exit 1
fi
# Clean: no drift. Warnings are expected (agents not dialed in) — stay silent.
exit 0
-342
View File
@@ -1,342 +0,0 @@
#!/usr/bin/env python3
"""identity-broker.py — Scope lifecycle over proxy providers.
Owns identity-state.json (gitignored runtime state: scope -> label
assignments, no secrets) and drives providers through it:
up <fp> resolve fingerprint, provision its scope
down <fp|unit> teardown scope network, drop assignment
cycle <fp> rotate to a fresh identity (bumps cycles)
exec <fp> -- <cmd> run a command inside the scope network
routes <fp> read-only route/tunnel status
status scopes with emails/labels (no key material)
bind --from <runs.json> --session <s> --fp <f>
attribute live runs (agent-manager --once
--json) to a scope after verifying the
session exists in the scan
v1 boundary: the broker scopes NETWORK identity only. It never sees,
stores, or prints key bytes — fingerprints and emails are the only
identifiers here. Short-lived brokered credential issuance is a
deferred stage; harnesses receive keys through existing means.
"""
from __future__ import annotations
import argparse
import importlib.util
import json
import sys
import time
from pathlib import Path
from typing import Any, Callable, Dict, List, Optional, Tuple
REPO_ROOT = Path(__file__).resolve().parent.parent
BIN_DIR = REPO_ROOT / "bin"
STATE_FILE = REPO_ROOT / "identity-state.json"
RunFn = Callable[..., Tuple[int, str]]
def _load(name: str, modname: str):
path = BIN_DIR / name
spec = importlib.util.spec_from_file_location(modname, path)
mod = importlib.util.module_from_spec(spec)
sys.modules[modname] = mod
spec.loader.exec_module(mod)
return mod
resolve_mod = _load("identity-resolve.py", "identity_resolve")
provider_mod = _load("identity-provider.py", "identity_provider")
PROVIDERS = {"warp": provider_mod.WarpProvider(),
"wireguard": provider_mod.GenericWireGuardProvider(),
"socks": provider_mod.SocksProxyProvider()}
class BrokerError(RuntimeError):
"""A broker operation failed (message is safe to show)."""
def load_state(path: str | Path = STATE_FILE) -> Dict[str, Any]:
"""Load broker state. Missing/corrupt -> empty (never raises)."""
try:
with open(path, "r") as f:
data = json.load(f)
if isinstance(data, dict):
data.setdefault("scopes", {})
data.setdefault("bindings", [])
return data
except Exception:
pass
return {"scopes": {}, "bindings": []}
def save_state(state: Dict[str, Any],
path: str | Path = STATE_FILE) -> None:
with open(path, "w") as f:
json.dump(state, f, indent=2, sort_keys=True)
def _provider(name: str = "warp"):
try:
prov = PROVIDERS[name]
except KeyError:
raise BrokerError("unknown provider %r (have: %s)"
% (name, ", ".join(sorted(PROVIDERS))))
if not prov.ready:
raise BrokerError("provider %r is boilerplate (not implemented); "
"warp is the live provider" % (name,))
return prov
def _scope_for_fp(fp: str, map_path=None) -> Dict[str, Any]:
scope = resolve_mod.resolve_scope(
resolve_mod.load_map(map_path or resolve_mod.MAP_FILE), fp)
if scope is None:
raise BrokerError("unknown fingerprint (not in identity-map.json)")
scope["slug"] = resolve_mod.scope_slug(scope)
return scope
def _unit_key(fp_or_unit: str, state: Dict[str, Any],
map_path=None) -> str:
"""Resolve CLI input (fp or scope unit) to a state scopes key."""
scopes = state.get("scopes", {})
if fp_or_unit in scopes:
return fp_or_unit
try:
scope = _scope_for_fp(fp_or_unit, map_path)
except BrokerError:
raise BrokerError("no scope for %r (unknown fingerprint, no "
"such unit)" % (fp_or_unit,))
if scope["unit"] not in scopes:
raise BrokerError("scope %s is not up" % (scope["unit"],))
return scope["unit"]
def op_up(fp: str, run: Optional[RunFn] = None, map_path=None,
state_path: str | Path = STATE_FILE,
provider_name: str = "warp") -> Dict[str, Any]:
"""Provision a scope's network. Idempotent (re-up returns existing)."""
scope = _scope_for_fp(fp, map_path)
state = load_state(state_path)
if scope["unit"] in state["scopes"]:
return {"ok": "exists", **state["scopes"][scope["unit"]]}
label = scope["slug"]
try:
res = _provider(provider_name).provision(label, run=run)
except provider_mod.ProviderError as e:
raise BrokerError(str(e))
rec = {"scope": scope["scope"], "email": scope["email"],
"origins": scope["origins"], "label": label,
"provider": provider_name, "netns": res.get("netns", ""),
"created": int(time.time()), "cycles": 0}
state["scopes"][scope["unit"]] = rec
save_state(state, state_path)
return {"ok": "true", **rec}
def op_down(fp_or_unit: str, run: Optional[RunFn] = None, map_path=None,
state_path: str | Path = STATE_FILE) -> Dict[str, Any]:
"""Teardown a scope's network and drop its assignment + bindings."""
state = load_state(state_path)
unit = _unit_key(fp_or_unit, state, map_path)
rec = state["scopes"][unit]
try:
_provider(rec.get("provider", "warp")).teardown(rec["label"],
run=run)
except provider_mod.ProviderError as e:
raise BrokerError(str(e))
del state["scopes"][unit]
state["bindings"] = [b for b in state.get("bindings", [])
if b.get("unit") != unit]
save_state(state, state_path)
return {"ok": "true", "unit": unit, "label": rec["label"]}
def op_cycle(fp: str, run: Optional[RunFn] = None, map_path=None,
state_path: str | Path = STATE_FILE) -> Dict[str, Any]:
"""Rotate a scope to a fresh identity (bumps the cycle count)."""
scope = _scope_for_fp(fp, map_path)
state = load_state(state_path)
if scope["unit"] not in state["scopes"]:
raise BrokerError("scope %s is not up (up it first)"
% (scope["unit"],))
rec = state["scopes"][scope["unit"]]
try:
_provider(rec.get("provider", "warp")).cycle(rec["label"],
run=run)
except provider_mod.ProviderError as e:
raise BrokerError(str(e))
rec["cycles"] = int(rec.get("cycles", 0)) + 1
save_state(state, state_path)
return {"ok": "true", "unit": scope["unit"], "label": rec["label"],
"cycles": rec["cycles"]}
def op_exec(fp: str, cmd: List[str], run: Optional[RunFn] = None,
map_path=None,
state_path: str | Path = STATE_FILE) -> Tuple[int, str]:
"""Run cmd inside the scope's network. Returns (rc, output)."""
scope = _scope_for_fp(fp, map_path)
state = load_state(state_path)
if scope["unit"] not in state["scopes"]:
raise BrokerError("scope %s is not up (up it first)"
% (scope["unit"],))
rec = state["scopes"][scope["unit"]]
try:
return _provider(rec.get("provider", "warp")).exec(
rec["label"], cmd, run=run)
except provider_mod.ProviderError as e:
raise BrokerError(str(e))
def op_routes(fp: str, run: Optional[RunFn] = None, map_path=None,
state_path: str | Path = STATE_FILE) -> Dict[str, Any]:
"""Read-only route/tunnel status for a scope."""
scope = _scope_for_fp(fp, map_path)
state = load_state(state_path)
if scope["unit"] not in state["scopes"]:
raise BrokerError("scope %s is not up (up it first)"
% (scope["unit"],))
rec = state["scopes"][scope["unit"]]
try:
info = _provider(rec.get("provider", "warp")).routes(
rec["label"], run=run)
except provider_mod.ProviderError as e:
raise BrokerError(str(e))
return {"unit": scope["unit"], "email": scope["email"], **info}
def op_status(run: Optional[RunFn] = None,
state_path: str | Path = STATE_FILE) -> Dict[str, Any]:
"""Scopes with live provider status. Emails/labels only."""
state = load_state(state_path)
scopes = []
for unit, rec in sorted(state.get("scopes", {}).items()):
try:
live = _provider(rec.get("provider", "warp")).status(
rec["label"], run=run)
except provider_mod.ProviderError as e:
live = {"conf": "?", "netns": "?", "egress": "n/a",
"error": str(e)}
scopes.append({"unit": unit, "email": rec.get("email", ""),
"scope": rec.get("scope", ""),
"label": rec.get("label", ""),
"provider": rec.get("provider", ""),
"cycles": rec.get("cycles", 0), **live})
return {"scopes": scopes, "bindings": state.get("bindings", [])}
def op_bind(runs_path: str, session: str, fp: str, map_path=None,
state_path: str | Path = STATE_FILE) -> Dict[str, Any]:
"""Attribute live runs to a scope, verifying against a scan.
runs_path is agent-manager.py --once --json output. Every run
whose session group contains `session` is bound to the fp's scope
(fp must resolve; the scope need not be up — binding is
attribution, not network). Sessions absent from the scan are
refused (never bind what we cannot observe).
"""
scope = _scope_for_fp(fp, map_path)
try:
with open(runs_path, "r") as f:
scan = json.load(f)
runs = scan.get("runs", [])
if not isinstance(runs, list):
raise ValueError("no runs list")
except Exception as e:
raise BrokerError("cannot read runs scan %s: %s" % (runs_path, e))
matched = []
for r in runs:
if not isinstance(r, dict):
continue
group = str(r.get("session", "")).split(",")
if session in group:
matched.append(r)
if not matched:
raise BrokerError("session %r not observed in %s (refusing to "
"bind unseen runs)" % (session, runs_path))
state = load_state(state_path)
now = int(time.time())
new = []
for r in matched:
rec = {"unit": scope["unit"], "email": scope["email"],
"device": r.get("device", ""), "type": r.get("type", ""),
"session": session, "pane": r.get("pane", ""),
"pid": r.get("pid", 0), "bound_at": now}
new.append(rec)
# Replace prior bindings for these exact runs (re-bind refreshes).
keys = {(b["device"], b.get("pane"), b.get("pid")) for b in new}
state["bindings"] = [b for b in state.get("bindings", [])
if (b.get("device"), b.get("pane"),
b.get("pid")) not in keys] + new
save_state(state, state_path)
return {"ok": "true", "unit": scope["unit"], "bound": len(new),
"runs": new}
def main(argv: Optional[List[str]] = None) -> int:
ap = argparse.ArgumentParser(prog="identity-broker.py")
ap.add_argument("--map", default=str(resolve_mod.MAP_FILE),
help="identity map (default: identity-map.json)")
ap.add_argument("--state", default=str(STATE_FILE),
help="broker state file (default: identity-state.json)")
sub = ap.add_subparsers(dest="cmd", required=True)
p = sub.add_parser("up", help="provision a scope network")
p.add_argument("fp")
p.add_argument("--provider", default="warp")
p = sub.add_parser("down", help="teardown a scope network")
p.add_argument("fp_or_unit")
p = sub.add_parser("cycle", help="rotate a scope identity")
p.add_argument("fp")
p = sub.add_parser("exec", help="run a command in a scope network")
p.add_argument("fp")
p.add_argument("exec_cmd", nargs=argparse.REMAINDER,
help="command (after --)")
p = sub.add_parser("routes", help="route/tunnel status for a scope")
p.add_argument("fp")
sub.add_parser("status", help="scopes + bindings (emails only)")
p = sub.add_parser("bind", help="attribute live runs to a scope")
p.add_argument("--from", dest="runs", required=True)
p.add_argument("--session", required=True)
p.add_argument("--fp", required=True)
args = ap.parse_args(argv)
mp, sp = args.map, args.state
try:
if args.cmd == "up":
print(json.dumps(op_up(args.fp, map_path=mp,
state_path=sp,
provider_name=args.provider),
indent=2))
elif args.cmd == "down":
print(json.dumps(op_down(args.fp_or_unit, map_path=mp,
state_path=sp), indent=2))
elif args.cmd == "cycle":
print(json.dumps(op_cycle(args.fp, map_path=mp,
state_path=sp), indent=2))
elif args.cmd == "exec":
cmd = [c for c in (args.exec_cmd or []) if c != "--"]
rc, out = op_exec(args.fp, cmd, map_path=mp,
state_path=sp)
sys.stdout.write(out + ("\n" if out else ""))
return rc
elif args.cmd == "routes":
print(json.dumps(op_routes(args.fp, map_path=mp,
state_path=sp), indent=2))
elif args.cmd == "status":
print(json.dumps(op_status(state_path=sp), indent=2))
elif args.cmd == "bind":
print(json.dumps(op_bind(args.runs, args.session, args.fp,
map_path=mp, state_path=sp),
indent=2))
return 0
except BrokerError as e:
print("error: %s" % e)
return 1
if __name__ == "__main__":
sys.exit(main())
-237
View File
@@ -1,237 +0,0 @@
#!/usr/bin/env python3
"""identity-provider.py — Proxy provider implementations.
A provider owns one network-identity substrate behind a fixed
interface: provision / teardown / cycle / exec / routes / status.
All subprocesses go through an injectable run function (same seam as
box-fleet-tui gather_*), so command shapes are unit-testable and no
test touches netns, sudo, or /etc/netvm.
Security boundaries (from the repo's own scripts):
- Warp identities generate via netvm-new-identity.sh, which the user
explicitly authorized operators to run (see netvm-provision-node.sh
header). Generation installs a root-0600 conf and prints nothing.
- This code NEVER reads /etc/netvm and never prints key material.
Confs are consumed only by root tools (wg setconf inside netns).
- CLI-facing output carries emails, labels, and fingerprints only.
"""
from __future__ import annotations
import os
import re
import subprocess
import sys
from pathlib import Path
from typing import Callable, Dict, List, Optional, Tuple
REPO_ROOT = Path(__file__).resolve().parent.parent
BIN_DIR = REPO_ROOT / "bin"
RunFn = Callable[..., Tuple[int, str]]
LABEL_RE = re.compile(r"^[a-z0-9][a-z0-9-]{0,22}$")
def _run(cmd: List[str], timeout: int = 120) -> Tuple[int, str]:
"""Run cmd, capture output. Returns (returncode, combined_output)."""
try:
r = subprocess.run(cmd, capture_output=True, text=True,
timeout=timeout)
return r.returncode, ((r.stdout or "") + (r.stderr or "")).strip()
except subprocess.TimeoutExpired:
return 124, "timed out after %ds: %s" % (timeout, " ".join(cmd))
except OSError as e:
return 127, str(e)
class ProviderError(RuntimeError):
"""A provider operation failed (message is safe to show)."""
def check_label(label: str) -> str:
"""Validate a netvm label. Returns it or raises ProviderError."""
if not LABEL_RE.match(label or ""):
raise ProviderError(
"invalid label %r: lowercase letters, digits, hyphens "
"(max 23 chars)" % (label,))
return label
class Provider:
"""Interface every proxy provider implements. Boilerplate subclasses
override these with real substrate calls; see WarpProvider."""
name = "base"
ready = False
def provision(self, label: str,
run: Optional[RunFn] = None) -> Dict[str, str]:
"""Create the network identity + bring it up. Idempotent."""
raise NotImplementedError
def teardown(self, label: str,
run: Optional[RunFn] = None) -> Dict[str, str]:
"""Bring the identity's network down (keeps the identity)."""
raise NotImplementedError
def cycle(self, label: str,
run: Optional[RunFn] = None) -> Dict[str, str]:
"""Rotate to a fresh identity (teardown + new identity + up)."""
raise NotImplementedError
def exec(self, label: str, cmd: List[str],
run: Optional[RunFn] = None) -> Tuple[int, str]:
"""Run cmd inside the identity's network. Returns (rc, output)."""
raise NotImplementedError
def routes(self, label: str,
run: Optional[RunFn] = None) -> Dict[str, str]:
"""Read-only route/tunnel status for the identity."""
raise NotImplementedError
def status(self, label: str,
run: Optional[RunFn] = None) -> Dict[str, str]:
"""Read-only liveness: conf present, netns up, egress IP."""
raise NotImplementedError
class WarpProvider(Provider):
"""Cloudflare Warp provider on the established warp-* structures.
Identity: /etc/netvm/<label>.conf via netvm-new-identity.sh
(operator-authorized). Network: warp-<label> netns via
netvm-node-up.sh / netvm-node-down.sh. Exec: netvm-exec.sh.
Scopes are NOT nodes: no chrome-box profile, no NODES.md entry.
"""
name = "warp"
ready = True
def _conf_exists(self, label: str, run: RunFn) -> bool:
rc, _ = run(["test", "-f", "/etc/netvm/%s.conf" % label],
timeout=10)
return rc == 0
def provision(self, label: str,
run: Optional[RunFn] = None) -> Dict[str, str]:
run = run or _run
check_label(label)
steps = []
if not self._conf_exists(label, run):
rc, out = run(["sudo", "-n", str(BIN_DIR / "netvm-new-identity.sh"),
label], timeout=300)
if rc != 0:
raise ProviderError("warp identity failed for %s: %s"
% (label, out[-200:]))
steps.append("identity=new")
else:
steps.append("identity=exists")
rc, out = run(["sudo", "-n", str(BIN_DIR / "netvm-node-up.sh"),
label], timeout=300)
if rc != 0:
raise ProviderError("netns up failed for %s: %s"
% (label, out[-200:]))
steps.append("netns=up")
return {"ok": "true", "label": label, "netns": "warp-" + label,
"steps": ",".join(steps)}
def teardown(self, label: str,
run: Optional[RunFn] = None) -> Dict[str, str]:
run = run or _run
check_label(label)
rc, out = run(["sudo", "-n", str(BIN_DIR / "netvm-node-down.sh"),
label], timeout=120)
if rc != 0:
raise ProviderError("netns down failed for %s: %s"
% (label, out[-200:]))
return {"ok": "true", "label": label, "netns": "warp-" + label}
def cycle(self, label: str,
run: Optional[RunFn] = None) -> Dict[str, str]:
"""Fresh warp identity: down + remove conf + provision.
Conf removal needs root on /etc/netvm; when denied, the old
identity is left intact (netns down) and the operator gets the
exact human step instead of a half-rotated state.
"""
run = run or _run
check_label(label)
self.teardown(label, run=run)
rc, out = run(["sudo", "-n", "rm", "-f",
"/etc/netvm/%s.conf" % label], timeout=30)
if rc != 0:
raise ProviderError(
"rotation paused for %s: cannot remove old identity "
"(%s). Human: sudo rm /etc/netvm/%s.conf, then cycle "
"again." % (label, out[-120:], label))
return self.provision(label, run=run)
def exec(self, label: str, cmd: List[str],
run: Optional[RunFn] = None) -> Tuple[int, str]:
run = run or _run
check_label(label)
if not cmd:
raise ProviderError("exec needs a command")
return run([str(BIN_DIR / "netvm-exec.sh"), label, "--"] + cmd,
timeout=120)
def routes(self, label: str,
run: Optional[RunFn] = None) -> Dict[str, str]:
run = run or _run
check_label(label)
netns = "warp-" + label
_, route_out = run(["sudo", "-n", "ip", "netns", "exec", netns,
"ip", "route"], timeout=30)
_, wg_out = run(["sudo", "-n", "ip", "netns", "exec", netns,
"wg", "show"], timeout=30)
return {"label": label, "netns": netns, "routes": route_out,
"wireguard": wg_out}
def status(self, label: str,
run: Optional[RunFn] = None) -> Dict[str, str]:
run = run or _run
check_label(label)
conf = self._conf_exists(label, run)
rc, out = run(["ip", "netns", "list"], timeout=10)
up = rc == 0 and ("warp-" + label) in out
egress = ""
if conf and up:
rc, eg = self.exec(label, ["curl", "-s", "--max-time", "8",
"https://api.ipify.org"], run=run)
egress = eg.strip().splitlines()[-1] if rc == 0 and eg.strip() \
else ""
return {"label": label, "conf": "yes" if conf else "no",
"netns": "up" if up else "down", "egress": egress or "n/a"}
class GenericWireGuardProvider(Provider):
"""BOILERPLATE: bring-your-own WireGuard confinement.
Intended structure: the operator supplies a wg conf out of band
(same root-0600 handling as Warp confs — never read here);
provision creates warp-<label> netns + veth/NAT exactly like
WarpProvider but consumes the supplied conf instead of a
Cloudflare-registered identity. Cycle swaps to the next supplied
conf. Implement when the first non-Cloudflare tunnel is needed.
"""
name = "wireguard"
class SocksProxyProvider(Provider):
"""BOILERPLATE: per-scope SOCKS5 forward, no netns.
Intended structure: provision opens a dedicated local forward
(ssh -D style) per scope label and records its port; exec runs
commands with ALL_PROXY scoped to that port instead of entering a
netns; cycle re-establishes the forward via a fresh egress.
Implement when a scope needs proxy semantics without tunnels.
"""
name = "socks"
if __name__ == "__main__":
print("identity-provider.py is a library (see identity-broker.py)")
sys.exit(2)
-186
View File
@@ -1,186 +0,0 @@
#!/usr/bin/env python3
"""identity-resolve.py — Pure identity resolution for the identity plane.
Reads identity-map.json (fingerprints only, never key material) and
resolves an API-key fingerprint to its network-identity scope:
api_key -> account_origin(s); one origin rolls scope UP to the
umbrella account, two or more keep scope DOWN at the key itself.
This module is pure + total (missing/corrupt map -> empty, unknown
fingerprint -> None). CLI output carries emails and fingerprints only;
key bytes never appear here — there is no code path that reads them
except `fp`, which hashes stdin and prints only the digest.
Usage:
identity-resolve.py fp < keyfile # print sha256: fingerprint
identity-resolve.py lookup <fingerprint> # print scope JSON
identity-resolve.py check # validate map schema
"""
from __future__ import annotations
import argparse
import hashlib
import json
import re
import sys
from pathlib import Path
from typing import Any, Dict, List, Optional
REPO_ROOT = Path(__file__).resolve().parent.parent
MAP_FILE = REPO_ROOT / "identity-map.json"
LABEL_RE = re.compile(r"^[a-z0-9][a-z0-9-]{0,22}$")
def fingerprint_hex(material: bytes) -> str:
"""sha256: fingerprint of raw key bytes."""
return "sha256:" + hashlib.sha256(material).hexdigest()
def load_map(path: str | Path = MAP_FILE) -> Dict[str, Any]:
"""Load the identity map. Missing/corrupt -> {"accounts": {}}."""
try:
with open(path, "r") as f:
data = json.load(f)
if isinstance(data, dict) and isinstance(
data.get("accounts"), dict):
return data
except Exception:
pass
return {"accounts": {}}
def find_key(map_data: Dict[str, Any],
fp: str) -> Optional[Dict[str, Any]]:
"""Locate a key record by fingerprint.
Returns {"email", "key"} or None. Top-level '_' entries ignored.
"""
if not fp:
return None
accounts = map_data.get("accounts")
if not isinstance(accounts, dict):
return None
for email, rec in accounts.items():
if not isinstance(rec, dict):
continue
keys = rec.get("keys")
if not isinstance(keys, list):
continue
for k in keys:
if isinstance(k, dict) and k.get("fp") == fp:
return {"email": email, "key": k}
return None
def resolve_scope(map_data: Dict[str, Any],
fp: str) -> Optional[Dict[str, Any]]:
"""Resolve a fingerprint to its scope unit.
Single origin -> {"scope": "account", "unit": email, ...}.
Multiple origins -> {"scope": "key", "unit": fp, ...}.
Unknown fingerprint -> None. Result carries emails + fingerprints
only (no key material exists anywhere in this module).
"""
found = find_key(map_data, fp)
if found is None:
return None
key = found["key"]
origins = key.get("origins")
if not isinstance(origins, list) or not origins:
return None
origins = [str(o) for o in origins]
if len(origins) == 1:
return {"scope": "account", "unit": origins[0],
"email": found["email"], "origins": origins,
"label": key.get("label", "")}
return {"scope": "key", "unit": fp, "email": found["email"],
"origins": origins, "label": key.get("label", "")}
def scope_slug(scope: Dict[str, Any]) -> str:
"""Deterministic netvm label for a scope (fits label validation).
Account scopes: id-<email-fragment>-<hash7>. Key scopes:
id-k-<fp-hex-prefix>. Always matches ^[a-z0-9][a-z0-9-]{0,22}$.
"""
unit = str(scope.get("unit", ""))
if scope.get("scope") == "key":
hexpart = re.sub(r"[^0-9a-f]", "", unit.lower())[:12] or "0"
return "id-k-%s" % hexpart
frag = re.sub(r"[^a-z0-9]+", "-", unit.lower()).strip("-")[:12]
frag = frag.strip("-") or "x"
tag = hashlib.sha256(unit.encode()).hexdigest()[:7]
return "id-%s-%s" % (frag, tag)
def check_map(map_data: Dict[str, Any]) -> List[str]:
"""Validate map schema. Returns a list of problem strings (empty OK)."""
problems: List[str] = []
accounts = map_data.get("accounts")
if not isinstance(accounts, dict):
return ["top-level 'accounts' must be an object"]
seen_fps: Dict[str, str] = {}
for email, rec in accounts.items():
if not isinstance(email, str) or "@" not in email:
problems.append("account key %r is not an email" % (email,))
if not isinstance(rec, dict) or not isinstance(
rec.get("keys"), list):
problems.append("account %r: 'keys' must be a list" % (email,))
continue
for i, k in enumerate(rec["keys"]):
where = "%s.keys[%d]" % (email, i)
if not isinstance(k, dict):
problems.append("%s: not an object" % where)
continue
fp = k.get("fp", "")
if not re.fullmatch(r"sha256:[0-9a-f]{64}", str(fp)):
problems.append("%s: bad fingerprint %r" % (where, fp))
elif fp in seen_fps:
problems.append("%s: fingerprint already listed under %s"
% (where, seen_fps[fp]))
else:
seen_fps[fp] = email
origins = k.get("origins")
if not isinstance(origins, list) or not origins or not all(
isinstance(o, str) and o for o in origins):
problems.append("%s: 'origins' must be a non-empty "
"string list" % where)
return problems
def main(argv: Optional[List[str]] = None) -> int:
ap = argparse.ArgumentParser(prog="identity-resolve.py")
ap.add_argument("--map", default=str(MAP_FILE),
help="identity map (default: identity-map.json)")
sub = ap.add_subparsers(dest="cmd", required=True)
sub.add_parser("fp", help="print sha256: fingerprint of stdin bytes")
p = sub.add_parser("lookup", help="resolve a fingerprint to scope JSON")
p.add_argument("fp")
sub.add_parser("check", help="validate the map schema")
args = ap.parse_args(argv)
if args.cmd == "fp":
sys.stdout.write(fingerprint_hex(sys.stdin.buffer.read()) + "\n")
return 0
if args.cmd == "lookup":
scope = resolve_scope(load_map(args.map), args.fp)
if scope is None:
print("unknown fingerprint (not in map)")
return 1
scope["slug"] = scope_slug(scope)
print(json.dumps(scope, indent=2))
return 0
problems = check_map(load_map(args.map))
if problems:
print("%s INVALID:" % args.map)
for prob in problems:
print(" - %s" % prob)
return 1
print("%s OK" % args.map)
return 0
if __name__ == "__main__":
sys.exit(main())
-296
View File
@@ -1,296 +0,0 @@
#!/usr/bin/env python3
"""Identity variance testing framework.
Sends gentle, controlled probes from each NetVM node's Warp identity and
measures per-identity response behavior so that future rate limits can be
attributed (per-identity vs per-IP vs time-based vs random).
Design goals:
- Very gentle traffic: 2 probes per node per cycle, default 10-min cycle.
- Never logs secrets: only sha256 hashes of WireGuard private keys.
- JSONL results log: one line per probe, machine-readable.
Usage:
identity-variance-test.py probe # run one probe cycle over all nodes
identity-variance-test.py analyze [--since HOURS] [--log PATH]
Probes per node:
1. https://www.cloudflare.com/cdn-cgi/trace (identity + egress IP Cloudflare sees)
2. https://muse.ai/ (production-relevant landing page)
Exit codes: 0 ok, 1 partial (some nodes failed), 2 fatal.
"""
import argparse
import hashlib
import json
import os
import re
import subprocess
import sys
import time
from datetime import datetime, timezone
DEFAULT_LOG = os.path.expanduser(
"~/Projects/NetVM/.state/identity-variance/variance.jsonl"
)
NODES = ["muse", "pip", "646", "opm", "def"]
PROBES = [
("cf-trace", "https://www.cloudflare.com/cdn-cgi/trace"),
("muse-landing", "https://muse.ai/"),
]
CURL_TIMEOUT = 15
def sh(cmd, timeout=30):
"""Run cmd (list) and return (rc, stdout, stderr)."""
try:
p = subprocess.run(
cmd, capture_output=True, text=True, timeout=timeout
)
return p.returncode, p.stdout, p.stderr
except subprocess.TimeoutExpired as e:
return 124, (e.stdout or ""), "timeout"
except FileNotFoundError:
return 127, "", "command not found"
def node_netns_exists(node):
rc, out, _ = sh(["sudo", "-n", "ip", "netns", "list"])
return rc == 0 and f"warp-{node}" in out
def identity_hash(node):
"""Return truncated sha256 of the node's WireGuard private key (never the key)."""
rc, out, _ = sh(["sudo", "-n", "cat", f"/etc/netvm/{node}.conf"])
if rc == 0:
for line in out.splitlines():
m = re.match(r"\s*PrivateKey\s*=\s*(\S+)", line)
if m:
return hashlib.sha256(m.group(1).encode()).hexdigest()[:16]
return None
def probe_node(node, target_name, url):
"""One probe from inside the node's netns. Returns dict."""
# -w fields: http_code, time_total, size_download, remote_ip
fmt = "%{http_code} %{time_total} %{size_download} %{remote_ip}"
cmd = [
"sudo", "-n", "ip", "netns", "exec", f"warp-{node}",
"curl", "-s", "-o", "/dev/null", "-m", str(CURL_TIMEOUT),
"-w", fmt, url,
]
started = datetime.now(timezone.utc)
rc, out, err = sh(cmd, timeout=CURL_TIMEOUT + 10)
ended = datetime.now(timezone.utc)
rec = {
"ts": started.isoformat(),
"node": node,
"probe": target_name,
"url": url,
"curl_rc": rc,
}
parts = out.strip().split()
if rc == 0 and len(parts) == 4:
try:
rec["http_status"] = int(parts[0])
except ValueError:
rec["http_status"] = None
try:
rec["response_ms"] = round(float(parts[1]) * 1000, 1)
except ValueError:
rec["response_ms"] = None
try:
rec["bytes"] = int(parts[2])
except ValueError:
rec["bytes"] = None
rec["remote_ip"] = parts[3] if parts[3] != "0.0.0.0" else None
else:
rec["http_status"] = None
rec["response_ms"] = None
rec["bytes"] = None
rec["remote_ip"] = None
rec["curl_error"] = (err or "curl failed").strip()[:200]
rec["rate_limited"] = rec["http_status"] in (429,)
rec["blocked"] = rec["http_status"] in (403,)
rec["ok"] = rec["http_status"] is not None and 200 <= rec["http_status"] < 400
rec["elapsed_wall_ms"] = round(
(ended - started).total_seconds() * 1000, 1
)
return rec
def egress_ip(node):
"""Best-effort egress IP as seen from inside the netns."""
rc, out, _ = sh([
"sudo", "-n", "ip", "netns", "exec", f"warp-{node}",
"curl", "-s", "-m", "10", "https://api.ipify.org",
], timeout=20)
if rc == 0 and re.fullmatch(r"[0-9a-fA-F.:]+", out.strip()):
return out.strip()
return None
def run_probe_cycle(log_path):
os.makedirs(os.path.dirname(log_path), exist_ok=True)
results = []
any_ok, any_fail = False, False
idhash = {}
for node in NODES:
if not node_netns_exists(node):
results.append({
"ts": datetime.now(timezone.utc).isoformat(),
"node": node, "probe": "node-skip",
"ok": False, "note": "netns warp-%s missing" % node,
})
any_fail = True
continue
idhash[node] = identity_hash(node)
for target_name, url in PROBES:
rec = probe_node(node, target_name, url)
rec["identity_hash"] = idhash[node]
results.append(rec)
if rec["ok"]:
any_ok = True
else:
any_fail = True
time.sleep(1) # gentle pacing between probes
# Attach egress IP per node (one lookup per node, cached per cycle).
egress = {}
for node in {r["node"] for r in results if r.get("probe") != "node-skip"}:
egress[node] = egress_ip(node)
for rec in results:
if rec.get("probe") != "node-skip":
rec["egress_ip"] = egress.get(rec["node"])
with open(log_path, "a") as f:
for rec in results:
f.write(json.dumps(rec) + "\n")
print(json.dumps({
"cycle_ts": datetime.now(timezone.utc).isoformat(),
"log": log_path,
"records": len(results),
"ok": sum(1 for r in results if r.get("ok")),
"failed": sum(1 for r in results if not r.get("ok")),
"nodes": sorted({r["node"] for r in results}),
"egress": egress,
"identities": idhash,
}, indent=2))
if not any_ok:
return 2
return 1 if any_fail else 0
def analyze(log_path, since_hours=None):
if not os.path.exists(log_path):
print("no log yet at %s" % log_path, file=sys.stderr)
return 2
cutoff = None
if since_hours:
cutoff = time.time() - since_hours * 3600
per_node = {}
rate_limit_events = []
identity_changes = {}
egress_by_node = {}
with open(log_path) as f:
for line in f:
line = line.strip()
if not line:
continue
try:
r = json.loads(line)
except json.JSONDecodeError:
continue
if r.get("probe") == "node-skip":
continue
try:
ts = datetime.fromisoformat(r["ts"]).timestamp()
except (ValueError, KeyError):
continue
if cutoff and ts < cutoff:
continue
node = r["node"]
st = per_node.setdefault(node, {
"probes": 0, "ok": 0, "ms": [], "429": 0, "403": 0,
"statuses": {}, "identities": set(), "probes_by_target": {},
})
st["probes"] += 1
if r.get("ok"):
st["ok"] += 1
if r.get("response_ms") is not None:
st["ms"].append(r["response_ms"])
if r.get("rate_limited"):
st["429"] += 1
rate_limit_events.append((r["ts"], node, r["probe"]))
if r.get("blocked"):
st["403"] += 1
s = r.get("http_status")
st["statuses"][str(s)] = st["statuses"].get(str(s), 0) + 1
if r.get("identity_hash"):
st["identities"].add(r["identity_hash"])
t = r.get("probe")
st["probes_by_target"][t] = st["probes_by_target"].get(t, 0) + 1
if r.get("egress_ip"):
egress_by_node.setdefault(node, set()).add(r["egress_ip"])
def pct(vals, p):
if not vals:
return None
s = sorted(vals)
return round(s[min(len(s) - 1, int(p / 100 * len(s)))], 1)
report = {"log": log_path, "nodes": {}}
for node in sorted(per_node):
st = per_node[node]
report["nodes"][node] = {
"probes": st["probes"],
"ok_rate": round(st["ok"] / st["probes"], 3) if st["probes"] else 0,
"latency_ms": {
"mean": round(sum(st["ms"]) / len(st["ms"]), 1) if st["ms"] else None,
"p50": pct(st["ms"], 50),
"p99": pct(st["ms"], 99),
},
"http_429": st["429"],
"http_403": st["403"],
"statuses": st["statuses"],
"identity_rotations": max(0, len(st["identities"]) - 1),
"egress_ips": sorted(egress_by_node.get(node, set())),
}
if len(st["identities"]) > 1:
identity_changes[node] = sorted(st["identities"])
# Divergence analysis: did all nodes see the same fate at the same time?
report["rate_limit_events"] = [
{"ts": ts, "node": n, "probe": p} for ts, n, p in rate_limit_events[-50:]
]
report["identity_changes"] = identity_changes
print(json.dumps(report, indent=2))
return 0
def main():
ap = argparse.ArgumentParser(description="Identity variance testing framework")
sub = ap.add_subparsers(dest="cmd", required=True)
p_probe = sub.add_parser("probe", help="run one probe cycle")
p_probe.add_argument("--log", default=DEFAULT_LOG)
p_an = sub.add_parser("analyze", help="summarize variance data")
p_an.add_argument("--log", default=DEFAULT_LOG)
p_an.add_argument("--since", type=float, default=None,
help="only include last N hours")
args = ap.parse_args()
if args.cmd == "probe":
sys.exit(run_probe_cycle(args.log))
sys.exit(analyze(args.log, args.since))
if __name__ == "__main__":
main()
-276
View File
@@ -1,276 +0,0 @@
#!/usr/bin/env python3
"""invite.py — Muse.ai invite codes + usage via the agent browsers.
Find: GET /api/hatch/invite via in-page fetch (primary; proven live),
Invite-button popover parse (DOM fallback).
Redeem: POST /api/hatch/invite-code/redeem via in-page fetch.
Usage: Settings > General DOM read (weekly reset, % used, additional
tokens) through the hatch_menu tree (dialog/mouse/tabs);
settings-nav primitives live there, this module keeps the
invite API + popover flows and the flat CLI parse.
Server eligibility (observed live): one redemption per account
(already_redeemed); 48h window from joining (window_expired); codes
carry limited uses (used_up) and can be revoked. Inviter rewards
accrue regardless of the inviter's own redemption state.
Caller errors (unknown node, malformed code) raise InviteError before
any CDP traffic. Transport/CDP failures return {"ok": False, ...}.
"""
import re
import time
from approvals import VALID_NODES, cdp_evaluate, get_cdp_ws
from hatch_menu import dialog as menu_dialog
from hatch_menu.mouse import close as _close
from hatch_menu.mouse import escape as _escape
from hatch_menu.tabs import general as general_tab
class InviteError(ValueError):
"""Caller error: unknown node or malformed code. Raised before CDP."""
CODE_RE = re.compile(r"^[A-Z0-9]{6}$")
CODE_REVEALED_RE = re.compile(r"Invite code revealed:\s*([A-Z0-9]{6})")
JS_INVITE_GET = """(async () => {
try {
const r = await fetch('/api/hatch/invite',
{method: 'GET', cache: 'no-store'});
return {status: r.status, data: await r.json()};
} catch (e) { return {error: String(e).slice(0, 200)}; }
})()"""
# %s is a validated [A-Z0-9]{6} code: injection-safe by construction.
JS_REDEEM_TMPL = """(async () => {
try {
const r = await fetch('/api/hatch/invite-code/redeem', {method: 'POST',
headers: {'Content-Type': 'application/json'},
body: JSON.stringify({code: '%s', supportsRedemptionStatus: true})});
return {status: r.status, ok: r.ok, data: await r.json()};
} catch (e) { return {error: String(e).slice(0, 200)}; }
})()"""
JS_INVITE_CLICK = ("(() => { const b = document.querySelector("
"'[data-testid=\"hatch-invite-friends-button\"]');"
" if (!b) return 'NO_BUTTON'; b.click();"
" return 'CLICKED'; })()")
JS_POPOVER_TEXT = ("(() => { const d = document.querySelector("
"'[data-slot=\"popover-content\"]');"
" return d ? d.innerText : null; })()")
def normalize_code(code):
"""Uppercase/strip a code; None unless 6-char A-Z0-9."""
if not isinstance(code, str):
return None
c = code.strip().upper()
return c if CODE_RE.fullmatch(c) else None
def _check_node(node):
if node not in VALID_NODES:
raise InviteError("unknown node: %r (valid: %s)"
% (node, ", ".join(VALID_NODES)))
def open_settings(ws):
"""Open the Settings dialog via the dock menu. True when open.
Public shim over hatch_menu.dialog.open_settings (single copy).
"""
return menu_dialog.open_settings(ws)
def click_settings_tab(ws, name):
"""Open a Settings dialog tab by visible name. True when open.
Public shim over hatch_menu.dialog.goto_tab (single copy).
"""
return menu_dialog.goto_tab(ws, name)
def _popover_code(ws):
"""Invite code via the main-chat popover (DOM fallback). None if absent."""
try:
if cdp_evaluate(ws, JS_INVITE_CLICK, timeout=5.0) != "CLICKED":
return None
time.sleep(1.0)
text = cdp_evaluate(ws, JS_POPOVER_TEXT, timeout=5.0) or ""
except Exception:
return None
finally:
_escape(ws)
m = CODE_REVEALED_RE.search(text)
return m.group(1) if m else None
def get_invite(node, timeout=8.0):
"""Per-agent invite state. API first, popover fallback for the code."""
_check_node(node)
try:
ws, _ = get_cdp_ws(node)
except Exception as e:
return {"ok": False, "node": node,
"error": "CDP unreachable: %s" % e}
try:
try:
res = cdp_evaluate(ws, JS_INVITE_GET, await_promise=True,
timeout=timeout)
except Exception as e:
res = {"error": "%s: %s" % (type(e).__name__, e)}
if isinstance(res, dict) and res.get("status") == 200 \
and isinstance(res.get("data"), dict) \
and res["data"].get("code"):
d = res["data"]
return {"ok": True, "node": node, "source": "api",
"code": d.get("code"),
"has_redeemed": d.get("has_redeemed_invite_code"),
"uses_remaining": d.get("uses_remaining"),
"use_count": d.get("use_count"),
"reward": d.get("reward")}
code = _popover_code(ws)
if code:
return {"ok": True, "node": node, "source": "dom",
"code": code, "has_redeemed": None,
"uses_remaining": None, "use_count": None,
"reward": None}
if isinstance(res, dict):
detail = res.get("error", "invite API failed")
else:
detail = "invite API failed"
return {"ok": False, "node": node,
"error": "invite lookup failed (%s); popover has no code"
% detail}
finally:
_close(ws)
def redeem_invite(node, code, timeout=12.0):
"""Redeem CODE on node. Returns ok / server reason + detail."""
c = normalize_code(code)
if c is None:
raise InviteError("malformed invite code: %r (want 6 chars A-Z0-9)"
% (code,))
_check_node(node)
js = JS_REDEEM_TMPL % c
try:
ws, _ = get_cdp_ws(node)
except Exception as e:
return {"ok": False, "node": node, "code": c,
"error": "CDP unreachable: %s" % e}
try:
try:
res = cdp_evaluate(ws, js, await_promise=True, timeout=timeout)
except Exception as e:
return {"ok": False, "node": node, "code": c,
"error": "%s: %s" % (type(e).__name__, e)}
if not isinstance(res, dict) or "data" not in res:
if isinstance(res, dict):
detail = res.get("error", "empty redeem response")
else:
detail = "empty redeem response"
return {"ok": False, "node": node, "code": c,
"reason": "unknown", "detail": detail}
data = res.get("data") or {}
if res.get("ok") and data.get("success"):
return {"ok": True, "node": node, "code": c,
"redemption_status": data.get("redemptionStatus"),
"detail": data.get("detail")}
return {"ok": False, "node": node, "code": c,
"reason": data.get("reason", "unknown"),
"detail": data.get("detail")}
finally:
_close(ws)
def parse_usage_text(text):
"""Flat usage parse (CLI contract) over the tree's General parser."""
nested = general_tab.parse_usage_text(text)
wk = (nested.get("weekly") if nested else None) or {}
add = (nested.get("additional") if nested else None) or {}
tokens = add.get("tokens_left")
return {
"weekly_reset": wk.get("resets_on"),
"weekly_used_pct": wk.get("pct_used"),
"additional_expires": "Never expires"
if add.get("never_expires") else None,
"additional_used_pct": add.get("pct_used"),
"additional_left": ("%s tokens left" % tokens)
if tokens else None,
"has_redeemed": bool(add.get("never_expires") or tokens
or "additional tokens" in (text or "").lower()),
}
def get_usage(node, timeout=8.0):
"""Usage limits for one agent via Settings General (DOM read)."""
_check_node(node)
try:
ws, _ = get_cdp_ws(node)
except Exception as e:
return {"ok": False, "node": node,
"error": "CDP unreachable: %s" % e}
try:
if not open_settings(ws):
return {"ok": False, "node": node,
"error": "settings dialog did not open"}
if not menu_dialog.goto_tab(ws, "General"):
return {"ok": False, "node": node,
"error": "General tab did not open"}
# Note: Usage stats and redeem field can load asynchronously in the React/Radix tree.
# Poll up to timeout seconds for usage or redeem markers.
text = None
deadline = time.time() + timeout
while time.time() < deadline:
t = menu_dialog.dialog_text(ws)
if t and ("Weekly limit" in t or "Additional tokens" in t or "Redeem invite code" in t or "tokens left" in t or "% used" in t):
text = t
break
time.sleep(0.4)
if not text:
text = menu_dialog.dialog_text(ws)
if not text:
return {"ok": False, "node": node,
"error": "empty settings dialog"}
out = {"ok": True, "node": node}
parsed = parse_usage_text(text)
stats_loaded = bool(
parsed.get("weekly_reset")
or parsed.get("weekly_used_pct") is not None
or parsed.get("has_redeemed")
or parsed.get("additional_left")
)
out.update(parsed)
out["stats_loaded"] = stats_loaded
if not stats_loaded:
out["note"] = "Usage stats did not render in Settings dialog"
return out
finally:
_escape(ws)
_close(ws)
def fleet_invite_status(nodes=None):
"""Invite state per node. Never raises; per-node error dicts."""
out = {}
for n in (nodes or list(VALID_NODES)):
try:
out[n] = get_invite(n)
except InviteError as e:
out[n] = {"ok": False, "node": n, "error": str(e)}
return out
def fleet_usage(nodes=None):
"""Usage limits per node. Never raises; per-node error dicts."""
out = {}
for n in (nodes or list(VALID_NODES)):
try:
out[n] = get_usage(n)
except InviteError as e:
out[n] = {"ok": False, "node": n, "error": str(e)}
return out
-790
View File
@@ -1,790 +0,0 @@
#!/usr/bin/env python3
"""invite_handler.py: Invite code discovery, inspection, and redemption handler.
Supports:
- Extracting agent invite codes via Main Chat DOM RPA & in-session API
- Redeeming invite codes via Settings Menu RPA & redemption endpoint
- Fleet-wide invite inventory, usage tracking, and agent salvage flows
"""
from __future__ import annotations
import argparse
import json
import re
import sys
import time
from dataclasses import asdict, dataclass
from typing import Any, Dict, List, Optional
try:
from approvals import cdp_evaluate, get_cdp_ws, get_node_pages
from settings_rpa import SettingsRPA, cdp_click_element_by_selector, cdp_send_escape
except ImportError:
import os
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
from approvals import cdp_evaluate, get_cdp_ws, get_node_pages
from settings_rpa import SettingsRPA, cdp_click_element_by_selector, cdp_send_escape
VALID_NODES = ["muse", "pip", "646", "opm", "def", "dev"]
@dataclass
class InviteCodeInfo:
node: str
code: str
uses_remaining: int
use_count: int
max_uses: int
has_redeemed: bool
invite_state: str
reward: Optional[Dict[str, Any]]
method_used: str
def to_dict(self) -> Dict[str, Any]:
return asdict(self)
@dataclass
class RedemptionResult:
target_node: str
code: str
success: bool
status: str
reason: Optional[str]
detail: Optional[str]
method_used: str
field_missing: bool = False
loopback_notified: bool = False
loopback_detail: Optional[str] = None
def to_dict(self) -> Dict[str, Any]:
return asdict(self)
def send_loopback_notice(
recipient: str,
target: str,
message: str,
sender: str = "super",
timeout: float = 15.0,
) -> bool:
"""Send a loopback notification DM via bin/dm.py to a sidechat without blocking or failing."""
try:
from pathlib import Path
bin_dir = Path(__file__).resolve().parent
dm_script = bin_dir / "dm.py"
if not dm_script.exists():
return False
import subprocess
cmd = [
sys.executable,
str(dm_script),
"send",
"--agent",
sender,
"--to",
recipient,
"--target",
target,
message,
]
if target == "main":
cmd.insert(-1, "--allow-main-chat")
res = subprocess.run(cmd, capture_output=True, text=True, timeout=timeout)
return res.returncode == 0
except Exception:
return False
class InviteHandler:
"""Handles invite codes discovery, inspection, and redemption for NetVM nodes."""
def __init__(self, node: str, timeout: float = 4.0):
self.node = node
self.timeout = timeout
self.ws = None
def _dispatch_loopback(self, res: RedemptionResult, agent: str, target: str) -> bool:
"""Post a loopback notification DM via box dm."""
msg = f"[BOX-INVITE-LOOPBACK] Node @{self.node} redeem status: {res.status}. Reason: {res.reason or 'unknown'}. Detail: {res.detail or '-'}"
ok = send_loopback_notice(recipient=agent, target=target, message=msg)
res.loopback_notified = ok
res.loopback_detail = f"Notified @{agent}/{target}" if ok else f"Failed to notify @{agent}/{target}"
return ok
def connect(self) -> InviteHandler:
if self.ws is None:
self.ws, _ = get_cdp_ws(self.node, timeout=self.timeout)
return self
def close(self) -> None:
if self.ws is not None:
try:
self.ws.close()
except Exception:
pass
self.ws = None
def __enter__(self) -> InviteHandler:
return self.connect()
def __exit__(self, exc_type, exc_val, exc_tb) -> None:
self.close()
def find_code_api(self) -> InviteCodeInfo:
"""Fetch invite code and metadata directly via in-session /api/hatch/invite query."""
self.connect()
js_query = """(async () => {
try {
const resp = await fetch('/api/hatch/invite', {method: 'GET', cache: 'no-store'});
if (!resp.ok) return {error: `HTTP ${resp.status}`};
return await resp.json();
} catch(e) {
return {error: e.toString()};
}
})()"""
res = cdp_evaluate(self.ws, js_query, await_promise=True, timeout=self.timeout)
if not res or "error" in res or "code" not in res:
err = res.get("error", "Invalid response") if res else "No response"
raise RuntimeError(f"Failed to fetch invite info for {self.node}: {err}")
return InviteCodeInfo(
node=self.node,
code=res.get("code", ""),
uses_remaining=res.get("uses_remaining", 0),
use_count=res.get("use_count", 0),
max_uses=res.get("max_uses", 30),
has_redeemed=bool(res.get("has_redeemed_invite_code", False)),
invite_state=res.get("invite_state", "UNKNOWN"),
reward=res.get("reward"),
method_used="api",
)
def find_code_dom(self) -> InviteCodeInfo:
"""Find invite code by navigating to Main Chat and opening the Invite popover."""
self.connect()
# 1. Ensure main chat home is active
with SettingsRPA(self.node, timeout=self.timeout) as rpa:
rpa.ensure_main_chat()
time.sleep(0.3)
# Dismiss any open popovers first
cdp_send_escape(self.ws)
time.sleep(0.2)
# 2. Click Invite button in main chat
clicked = cdp_click_element_by_selector(self.ws, '[data-testid="hatch-invite-friends-button"]', timeout=self.timeout)
if not clicked:
# Fallback to button by text
js_click = """(() => {
const btn = Array.from(document.querySelectorAll('button')).find(b => (b.innerText||'').trim() === 'Invite');
if (btn) { btn.click(); return true; }
return false;
})()"""
clicked = cdp_evaluate(self.ws, js_click)
if not clicked:
raise RuntimeError(f"Could not click Invite button on {self.node}")
time.sleep(0.8)
# 3. Read popover contents
js_read_popover = """(() => {
const pop = document.querySelector('[data-slot="popover-content"]');
if (!pop) return {found: false};
// Look for revealed code span
const codeSpan = pop.querySelector('[data-pel-impression="invite_code_revealed_impression"]');
let code = codeSpan ? codeSpan.innerText.trim() : null;
// Regex fallback on text
const fullText = pop.innerText || '';
if (!code) {
const m = fullText.match(/Invite code revealed:\\s*([A-Z0-9]{6})/);
if (m) code = m[1];
}
if (!code) {
const m2 = fullText.match(/\\b([A-Z0-9]{6})\\b/);
if (m2) code = m2[1];
}
let usesLeft = 30;
const mUses = fullText.match(/(\\d+)\\s+uses\\s+left/);
if (mUses) usesLeft = parseInt(mUses[1], 10);
return {
found: true,
code: code,
uses_left: usesLeft,
text: fullText
};
})()"""
pop_res = cdp_evaluate(self.ws, js_read_popover)
cdp_send_escape(self.ws)
if not pop_res or not pop_res.get("found") or not pop_res.get("code"):
# Fall back to API if popover extraction failed
return self.find_code_api()
code = pop_res.get("code")
uses_left = pop_res.get("uses_left", 30)
# Query API for additional metadata
try:
api_info = self.find_code_api()
api_info.method_used = "dom+api"
return api_info
except Exception:
return InviteCodeInfo(
node=self.node,
code=code,
uses_remaining=uses_left,
use_count=30 - uses_left,
max_uses=30,
has_redeemed=True,
invite_state="ELIGIBLE",
reward=None,
method_used="dom",
)
def find_code(self, method: str = "auto") -> InviteCodeInfo:
"""Find invite code using specified method ('auto', 'api', or 'dom')."""
if method == "api":
return self.find_code_api()
elif method == "dom":
return self.find_code_dom()
else:
# Auto: Try API first, fallback to DOM
try:
return self.find_code_api()
except Exception:
return self.find_code_dom()
def redeem_code_api(self, code: str) -> RedemptionResult:
"""Redeem invite code using in-session /api/hatch/invite-code/redeem endpoint."""
self.connect()
clean_code = code.strip().upper()
js_redeem = f"""(async () => {{
try {{
const resp = await fetch('/api/hatch/invite-code/redeem', {{
method: 'POST',
headers: {{'Content-Type': 'application/json'}},
body: JSON.stringify({{code: {json.dumps(clean_code)}, supportsRedemptionStatus: true}})
}});
const data = await resp.json();
return {{status_code: resp.status, ok: resp.ok, data: data}};
}} catch(e) {{
return {{error: e.toString()}};
}}
}})()"""
res = cdp_evaluate(self.ws, js_redeem, await_promise=True, timeout=self.timeout)
if not res or "error" in res:
err = res.get("error", "No response") if res else "Timeout"
return RedemptionResult(
target_node=self.node,
code=clean_code,
success=False,
status="error",
reason="transport_failure",
detail=err,
method_used="api",
)
data = res.get("data", {})
ok = res.get("ok", False) and data.get("success", False)
status = data.get("redemptionStatus") or ("redeemed" if ok else "failed")
reason = data.get("reason")
detail = data.get("detail")
return RedemptionResult(
target_node=self.node,
code=clean_code,
success=ok,
status=status,
reason=reason,
detail=detail,
method_used="api",
)
def redeem_code_dom(
self,
code: str,
notify_target: Optional[str] = None,
notify_agent: Optional[str] = None,
) -> RedemptionResult:
"""Redeem invite code using Settings Menu RPA -> Settings -> Redeem Invite Code dialog.
Gracefully handles missing entrypoint rows, missing input fields, and async loading races.
Falls back to in-page API automatically if DOM fields are absent, and dispatches DM loopback if requested."""
self.connect()
clean_code = code.strip().upper()
with SettingsRPA(self.node, timeout=self.timeout) as rpa:
# 1. Open Settings dialog
opened = rpa.open_settings_dialog()
if not opened:
# If dialog failed to open, try API fallback directly
api_res = self.redeem_code_api(clean_code)
if api_res.success:
api_res.detail = f"Settings dialog could not open; redemption completed via API fallback ({api_res.detail or ''})".strip()
return api_res
res = RedemptionResult(
target_node=self.node,
code=clean_code,
success=False,
status="error",
reason="settings_open_failed",
detail=f"Could not open Settings dialog via RPA on @{self.node}. API fallback: {api_res.detail or api_res.reason or 'failed'}",
method_used="dom",
field_missing=True,
)
if notify_target and notify_agent:
self._dispatch_loopback(res, notify_agent, notify_target)
return res
rpa.select_tab("General")
time.sleep(0.3)
# 2. Check for "Redeem invite code" entry (poll up to 2.5s for React rendering)
js_find_and_click = """(() => {
const dialog = document.querySelector('[role="dialog"]');
if (!dialog) return {found: false, dialog_present: false};
const items = Array.from(dialog.querySelectorAll('button, div, span'));
const redeemItem = items.find(el => (el.innerText || '').trim() === 'Redeem invite code');
if (redeemItem) {
redeemItem.click();
return {found: true, dialog_present: true};
}
return {
found: false,
dialog_present: true,
text: dialog.innerText || '',
has_additional: (dialog.innerText || '').includes('Additional tokens')
};
})()"""
click_res = None
deadline = time.time() + 2.5
while time.time() < deadline:
click_res = cdp_evaluate(self.ws, js_find_and_click)
if click_res and click_res.get("found"):
break
time.sleep(0.3)
if not click_res or not click_res.get("found"):
# "Redeem invite code" field is missing in General settings!
rpa.close_settings_dialog()
# Attempt automatic API fallback first
api_res = self.redeem_code_api(clean_code)
if api_res.success:
api_res.detail = f"Redeem field was missing in Settings DOM; redeemed successfully via API fallback! ({api_res.detail or ''})".strip()
return api_res
# Diagnose why field is missing
diag_text = click_res.get("text", "") if click_res else ""
has_extra = click_res.get("has_additional", False) if click_res else False
is_already = has_extra or "Additional tokens" in diag_text
if not is_already:
try:
api_check = self.find_code_api()
is_already = bool(api_check.has_redeemed)
except Exception:
pass
if is_already:
status = "already_redeemed"
reason = "already_redeemed"
detail = f"Node @{self.node} has already redeemed an invite code (entrypoint hidden by active Additional tokens ticker)."
elif api_res.reason in ["window_expired", "already_redeemed", "invalid", "used_up"]:
status = api_res.status
reason = api_res.reason
detail = f"Redeem invite code field not present in Settings on @{self.node}: server reports {api_res.reason} ({api_res.detail or ''})."
else:
status = "entrypoint_not_found"
reason = "field_missing"
detail = f"Redeem invite code field missing in General settings on @{self.node} (account may be past 48h onboarding window or already redeemed)."
res = RedemptionResult(
target_node=self.node,
code=clean_code,
success=False,
status=status,
reason=reason,
detail=detail,
method_used="dom",
field_missing=True,
)
if notify_target and notify_agent:
self._dispatch_loopback(res, notify_agent, notify_target)
return res
# 3. Handle HatchInviteRedemptionDialog (sub-dialog opened by clicking Redeem)
# Poll up to 2.5s for input box to mount
js_input_code = f"""(() => {{
// Look for dialog titled "Redeem a code" or input with label
const input = document.querySelector('input[aria-label="Invite code"]') ||
document.querySelector('[role="dialog"] input[type="text"]');
if (!input) return {{found_input: false}};
input.focus();
input.value = {json.dumps(clean_code)};
input.dispatchEvent(new Event('input', {{bubbles: true}}));
input.dispatchEvent(new Event('change', {{bubbles: true}}));
input.dispatchEvent(new KeyboardEvent('keydown', {{key: 'Enter', code: 'Enter', keyCode: 13, which: 13, bubbles: true}}));
return {{found_input: true}};
}})()"""
input_res = None
deadline = time.time() + 2.5
while time.time() < deadline:
input_res = cdp_evaluate(self.ws, js_input_code)
if input_res and input_res.get("found_input"):
break
time.sleep(0.3)
if not input_res or not input_res.get("found_input"):
# Subdialog opened or clicked, but input box is absent!
cdp_send_escape(self.ws)
rpa.close_settings_dialog()
# Automatic API fallback
api_res = self.redeem_code_api(clean_code)
if api_res.success:
api_res.detail = f"Redemption input box was missing in dialog; redeemed successfully via API fallback! ({api_res.detail or ''})".strip()
return api_res
res = RedemptionResult(
target_node=self.node,
code=clean_code,
success=False,
status=api_res.status or "dom_input_missing",
reason=api_res.reason or "input_field_missing",
detail=f"Invite code input field was not found in redemption dialog on @{self.node}. API fallback: {api_res.detail or api_res.reason or 'failed'}.",
method_used="dom",
field_missing=True,
)
if notify_target and notify_agent:
self._dispatch_loopback(res, notify_agent, notify_target)
return res
# Input submitted; poll for outcome text
time.sleep(1.2)
js_check_outcome = """(() => {
const dialog = document.querySelector('[role="dialog"]');
if (!dialog) return {found: false};
const text = dialog.innerText || '';
if (text.includes('Congratulations') || text.includes('successful')) {
return {success: true, status: 'redeemed', detail: text};
}
if (text.includes('already redeemed')) {
return {success: false, status: 'already_redeemed', detail: text};
}
if (text.includes('no longer valid') || text.includes('invalid')) {
return {success: false, status: 'invalid', detail: text};
}
if (text.includes('used up')) {
return {success: false, status: 'used_up', detail: text};
}
if (text.includes('passed') || text.includes('expired')) {
return {success: false, status: 'window_expired', detail: text};
}
return {status: 'unknown', detail: text};
})()"""
outcome = None
deadline = time.time() + 2.5
while time.time() < deadline:
outcome = cdp_evaluate(self.ws, js_check_outcome)
if outcome and outcome.get("status") != "unknown":
break
time.sleep(0.3)
rpa.close_settings_dialog()
if outcome and outcome.get("success"):
return RedemptionResult(
target_node=self.node,
code=clean_code,
success=True,
status="redeemed",
reason=None,
detail="Redemption successful via Settings RPA",
method_used="dom",
)
elif outcome and outcome.get("status") in ["already_redeemed", "invalid", "used_up", "window_expired"]:
res = RedemptionResult(
target_node=self.node,
code=clean_code,
success=False,
status=outcome.get("status"),
reason=outcome.get("status"),
detail=outcome.get("detail"),
method_used="dom",
)
if notify_target and notify_agent:
self._dispatch_loopback(res, notify_agent, notify_target)
return res
# Fallback to API check if DOM did not confirm status
api_res = self.redeem_code_api(clean_code)
if not api_res.success and notify_target and notify_agent:
self._dispatch_loopback(api_res, notify_agent, notify_target)
return api_res
def redeem_code(
self,
code: str,
method: str = "auto",
notify_target: Optional[str] = None,
notify_agent: Optional[str] = None,
) -> RedemptionResult:
"""Redeem invite code using specified method ('auto', 'dom', or 'api')."""
if method == "dom":
return self.redeem_code_dom(code, notify_target=notify_target, notify_agent=notify_agent)
elif method == "api":
res = self.redeem_code_api(code)
if not res.success and notify_target and notify_agent:
self._dispatch_loopback(res, notify_agent, notify_target)
return res
else:
# Auto: Validate and attempt via API for reliability. Fallback to DOM on transport error.
res = self.redeem_code_api(code)
if not res.success and res.reason == "transport_failure":
res = self.redeem_code_dom(code, notify_target=notify_target, notify_agent=notify_agent)
elif not res.success and notify_target and notify_agent:
self._dispatch_loopback(res, notify_agent, notify_target)
return res
def scan_fleet_invites(nodes: Optional[List[str]] = None) -> List[Dict[str, Any]]:
"""Scan fleet nodes and return invite code details for each active node."""
target_nodes = nodes or ["muse", "pip", "646", "opm"]
results = []
for node in target_nodes:
try:
handler = InviteHandler(node)
with handler:
info = handler.find_code(method="auto")
results.append(info.to_dict())
except Exception as e:
results.append({
"node": node,
"code": None,
"uses_remaining": 0,
"use_count": 0,
"max_uses": 30,
"has_redeemed": None,
"invite_state": "UNREACHABLE",
"reward": None,
"method_used": "error",
"error": str(e),
})
return results
def scan_fleet_usage(nodes: Optional[List[str]] = None) -> List[Dict[str, Any]]:
"""Scan fleet nodes and return token usage for each active node."""
target_nodes = nodes or ["muse", "pip", "646", "opm"]
results = []
for node in target_nodes:
try:
with SettingsRPA(node) as rpa:
usage = rpa.read_usage()
results.append(usage.to_dict())
except Exception as e:
results.append({
"node": node,
"plan": "Unknown",
"weekly_reset_text": "Unreachable",
"weekly_percent_used": 0,
"extra_tokens_status": "Unknown",
"extra_percent_used": 0,
"extra_tokens_remaining": "Unknown",
"is_blocked": False,
"bars": [],
"error": str(e),
})
return results
def salvage_blocked_node(
blocked_node: str = "646",
helper_node: Optional[str] = None,
notify_target: Optional[str] = None,
notify_agent: Optional[str] = None,
) -> Dict[str, Any]:
"""Salvage an out-of-tokens node by identifying its code and redeeming it on an eligible peer."""
# 1. Fetch blocked node invite code
with InviteHandler(blocked_node) as h_blocked:
blocked_info = h_blocked.find_code()
code_to_redeem = blocked_info.code
if not code_to_redeem:
return {
"success": False,
"blocked_node": blocked_node,
"error": f"Could not find invite code for blocked node {blocked_node}",
}
# 2. Check candidate helper nodes
candidates = [helper_node] if helper_node else [n for n in VALID_NODES if n != blocked_node]
eligible_peer = None
for peer in candidates:
try:
with InviteHandler(peer) as h_peer:
peer_info = h_peer.find_code()
if not peer_info.has_redeemed:
eligible_peer = peer
break
except Exception:
continue
if not eligible_peer:
res = {
"success": False,
"blocked_node": blocked_node,
"invite_code": code_to_redeem,
"error": "No existing fleet peer is currently eligible (all active peers have already redeemed an invite code). An onboarding agent or fresh client profile must redeem this code.",
"code_to_redeem": code_to_redeem,
"share_instruction": f"Redeem code '{code_to_redeem}' on a newly provisioned agent to credit 1 billion tokens to {blocked_node}.",
"field_missing": True,
"loopback_notified": False,
}
if notify_target and notify_agent:
msg = f"[SALVAGE-NOTICE] Node @{blocked_node} is blocked (code: {code_to_redeem}), but no eligible peer is available. Fresh onboarding required."
res["loopback_notified"] = send_loopback_notice(notify_agent, notify_target, msg)
return res
# 3. Redeem on eligible peer
with InviteHandler(eligible_peer) as h_peer:
redemption = h_peer.redeem_code(
code_to_redeem,
notify_target=notify_target,
notify_agent=notify_agent,
)
return {
"success": redemption.success,
"blocked_node": blocked_node,
"helper_node": eligible_peer,
"code_redeemed": code_to_redeem,
"redemption_result": redemption.to_dict(),
"field_missing": redemption.field_missing,
"loopback_notified": redemption.loopback_notified,
}
def main() -> None:
parser = argparse.ArgumentParser(description="NetVM Invite Code Handler")
subparsers = parser.add_subparsers(dest="command")
p_find = subparsers.add_parser("find", help="Find invite code for an agent")
p_find.add_argument("node", help="Node name (e.g. 646, pip, muse, opm)")
p_find.add_argument("--method", choices=["auto", "api", "dom"], default="auto")
p_find.add_argument("--json", action="store_true")
p_redeem = subparsers.add_parser("redeem", help="Redeem an invite code on a target agent")
p_redeem.add_argument("node", help="Target node to redeem the code on")
p_redeem.add_argument("code", help="6-character invite code")
p_redeem.add_argument("--method", choices=["auto", "api", "dom"], default="auto")
p_redeem.add_argument("--notify-target", default=None, help="Sidechat to notify on loopback")
p_redeem.add_argument("--notify-agent", default=None, help="Agent to notify on loopback")
p_redeem.add_argument("--json", action="store_true")
p_list = subparsers.add_parser("list", help="List invite codes across fleet")
p_list.add_argument("--json", action="store_true")
p_salvage = subparsers.add_parser("salvage", help="Salvage a blocked agent (e.g. 646)")
p_salvage.add_argument("node", default="646", nargs="?", help="Blocked node (default: 646)")
p_salvage.add_argument("--helper", help="Specific helper node to redeem code")
p_salvage.add_argument("--notify-target", default=None, help="Sidechat to notify on loopback")
p_salvage.add_argument("--notify-agent", default=None, help="Agent to notify on loopback")
p_salvage.add_argument("--json", action="store_true")
args = parser.parse_args()
if args.command == "find":
with InviteHandler(args.node) as h:
info = h.find_code(method=args.method)
if args.json:
print(json.dumps(info.to_dict(), indent=2))
else:
print(f"=== Agent {info.node} Invite Code ===")
print(f" Code: {info.code}")
print(f" Uses Remaining: {info.uses_remaining} / {info.max_uses}")
print(f" Has Redeemed?: {'Yes' if info.has_redeemed else 'No'}")
print(f" Method: {info.method_used}")
if info.reward:
print(f" Reward: {info.reward.get('title')}")
elif args.command == "redeem":
with InviteHandler(args.node) as h:
res = h.redeem_code(
args.code,
method=args.method,
notify_target=getattr(args, "notify_target", None),
notify_agent=getattr(args, "notify_agent", None),
)
if args.json:
print(json.dumps(res.to_dict(), indent=2))
else:
status_icon = "✔" if res.success else "✖"
print(f"[{status_icon}] Redemption on {res.target_node} for code {res.code}:")
print(f" Success: {res.success}")
print(f" Status: {res.status}")
if res.detail:
print(f" Detail: {res.detail}")
if res.loopback_notified:
print(f" Loopback:{res.loopback_detail}")
elif args.command == "list":
fleet = scan_fleet_invites()
if args.json:
print(json.dumps(fleet, indent=2))
else:
print("\n=== NETVM FLEET INVITE CODES ===")
print(f" {'NODE':<8} {'INVITE CODE':<14} {'USES LEFT':<12} {'REDEEMED?':<12} {'REWARD / NOTE':<25}")
print(f" {'────':<8} {'───────────':<14} {'─────────':<12} {'─────────':<12} {'─────────────':<25}")
for row in fleet:
node = row.get("node", "")
code = row.get("code") or "N/A"
uses = f"{row.get('uses_remaining', 0)}/{row.get('max_uses', 30)}"
redeemed = "Yes" if row.get("has_redeemed") else "No"
reward = row.get("reward", {})
reward_str = reward.get("title", "-") if reward else "-"
print(f" {node:<8} {code:<14} {uses:<12} {redeemed:<12} {reward_str:<25}")
print()
elif args.command == "salvage":
salvage_res = salvage_blocked_node(
args.node,
helper_node=args.helper,
notify_target=getattr(args, "notify_target", None),
notify_agent=getattr(args, "notify_agent", None),
)
if args.json:
print(json.dumps(salvage_res, indent=2))
else:
print(f"\n=== SALVAGE REPORT FOR {args.node.upper()} ===")
print(f" Invite Code to Credit: {salvage_res.get('invite_code') or salvage_res.get('code_redeemed')}")
if salvage_res.get("success"):
print(f" Status: SUCCESS! Redeemed on {salvage_res.get('helper_node')}")
else:
print(f" Status: {salvage_res.get('error')}")
if salvage_res.get("share_instruction"):
print(f" Next Step: {salvage_res.get('share_instruction')}")
if salvage_res.get("loopback_notified"):
print(" Loopback: Notice sent to requesting target.")
print()
else:
parser.print_help()
if __name__ == "__main__":
main()
-602
View File
@@ -1,602 +0,0 @@
#!/usr/bin/env python3
"""
job-dispatch.py: Dispatch a job by sending a DM to an agent.
Usage:
job-dispatch.py <job_name> [--dry-run]
Reads /home/super/Projects/NetVM/jobs/<job_name>.json,
renders the prompt template, sends DM via dm.py, logs to job-log.jsonl.
Part of the JOB system (see docs/JOB-SPEC.md).
Sidechat-first policy (2026-10-04): a job that resolves to target "main"
without explicit opt-in fails loudly instead of saturating main threads.
Opt in via job JSON "allow_main_chat": true, or --allow-main-chat.
"""
import sys
import re
import os
import json
import subprocess
import uuid
import argparse
from datetime import datetime, timezone
from pathlib import Path
# Paths
NETVM_ROOT = Path("/home/super/Projects/NetVM")
sys.path.insert(0, str(NETVM_ROOT / "bin"))
try:
import pipeline_engine
HAS_PIPELINE = True
except ImportError:
HAS_PIPELINE = False
JOBS_DIR = NETVM_ROOT / "jobs"
DM_PY = NETVM_ROOT / "bin" / "dm.py"
CHAT_API = NETVM_ROOT / "bin" / "muse-chat-api.py"
NETVM_EXEC = "/home/super/Projects/NetVM/bin/netvm-exec.sh"
JOB_LOG = NETVM_ROOT / "job-log.jsonl"
SIDECHAT_STATE = NETVM_ROOT / "job-sidechats.json"
def load_sidechat_state():
if SIDECHAT_STATE.exists():
try:
return json.loads(SIDECHAT_STATE.read_text())
except Exception:
return {}
return {}
def save_sidechat_state(state):
tmp = SIDECHAT_STATE.with_suffix(".tmp")
tmp.write_text(json.dumps(state, indent=2))
tmp.replace(SIDECHAT_STATE)
def extract_uuid(url):
m = re.search(r"/thread/([0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12})", url or "")
return m.group(1) if m else None
def get_current_url(agent="opm"):
try:
cmd = [NETVM_EXEC, agent, "--", "python3", str(CHAT_API),
"--account", agent, "url"]
r = subprocess.run(cmd, capture_output=True, text=True, timeout=30)
if r.returncode == 0:
return r.stdout.strip()
except Exception:
pass
return None
def use_sidechat_uuid(agent, thread_uuid, dry_run=False):
if dry_run:
return False
cmd = [NETVM_EXEC, agent, "--", "python3", str(CHAT_API),
"--account", agent, "sidechat", "use", thread_uuid]
try:
r = subprocess.run(cmd, capture_output=True, text=True, timeout=30)
return r.returncode == 0
except Exception:
return False
# Add bin to path for rate_limiter
sys.path.insert(0, str(NETVM_ROOT / "bin"))
try:
from rate_limiter import rate_limit_wait
HAS_RATE_LIMITER = True
except ImportError:
HAS_RATE_LIMITER = False
def log_event(event_type, data):
"""Append event to job-log.jsonl"""
entry = {
"ts": datetime.now(timezone.utc).isoformat(),
"type": event_type,
**data
}
with open(JOB_LOG, "a") as f:
f.write(json.dumps(entry) + "\n")
def load_job(job_name):
"""Load job JSON definition"""
# Try .json first, then .yaml (for compatibility)
job_file = JOBS_DIR / f"{job_name}.json"
if not job_file.exists():
archived_file = JOBS_DIR / "archive" / f"{job_name}.json"
if archived_file.exists():
print(f"Error: Job '{job_name}' is archived at {archived_file}. Unarchive before dispatch (e.g. 'box job unarchive {job_name}').", file=sys.stderr)
sys.exit(1)
job_file = JOBS_DIR / f"{job_name}.yaml"
if not job_file.exists():
print(f"Error: Job '{job_name}' not found in {JOBS_DIR}", file=sys.stderr)
sys.exit(1)
print(f"Warning: YAML not supported (no PyYAML). Convert {job_file} to JSON.", file=sys.stderr)
sys.exit(1)
with open(job_file) as f:
return json.load(f)
def render_prompt(template, variables):
"""Render prompt template with variables"""
# Simple {var} substitution
result = template
for key, value in variables.items():
result = result.replace(f"{{{key}}}", str(value))
return result
# ---- follow-up tracking (DM follow-up system integration) ----------------
# Jobs opt in via a "followup" block in the job JSON:
#
# "followup": {
# "expect_reply": true, # required: enables tracking
# "timeout": "1h", # duration ("30s","15m","2h","1d") or seconds
# # int; default "1h" (3600s)
# "nudges": 2, # 0..10, default 2
# "escalate": "opm", # identity string, default "opm"
# "route": "646-pip-coord" # optional route_id
# }
#
# The dispatcher translates this into dm.py --tag flags using the canonical
# vocabulary (box-threads/DEPLOY-DECISIONS.md). dm.py strips the tags from
# delivered text and creates a dm_followup request-store record after
# SENT+VERIFIED. Jobs without a followup block behave exactly as today.
#
# LIMITATIONS (v1):
# - The heartbeat job NEVER gets follow-ups (loopback health check).
# Hardcoded guard below; a followup block on heartbeat is ignored loudly.
# - Sidechat sends (muse-chat-api.py direct path) do not go through dm.py,
# so --tag flags cannot attach. v2 needs a record-creation path that does
# not send (e.g. POST /api/box/followups, or a bl->VM queue; bl cannot
# currently SSH to the VM). The dispatcher logs a warning when a
# sidechat-targeted job has followup enabled.
HEARTBEAT_JOB_NAME = "heartbeat"
def parse_followup_duration(value):
"""Parse a followup timeout into seconds. Accepts int (seconds) or
strings like '30s', '15m', '2h', '1d'. Returns int seconds.
Raises ValueError on bad input."""
if isinstance(value, int) and not isinstance(value, bool):
s = value
elif isinstance(value, str):
m = re.fullmatch(r"(\d+)\s*([smhd])?", value.strip().lower())
if not m:
raise ValueError("bad duration %r" % (value,))
n = int(m.group(1))
unit = m.group(2) or "s"
s = n * {"s": 1, "m": 60, "h": 3600, "d": 86400}[unit]
else:
raise ValueError("timeout must be int seconds or duration string")
if not 60 <= s <= 604800:
raise ValueError("timeout must be 60..604800s (1m..7d), got %d" % s)
return s
def build_followup_tags(followup):
"""Translate a job's followup block into dm.py --tag arguments.
Returns a flat list like ['--tag', 'reply:timeout=3600', ...].
Returns [] if followup is falsy or expect_reply is not true.
Raises ValueError on invalid config (caller logs a warning and sends
the DM untagged -- the job itself must never fail over this)."""
if not followup or not followup.get("expect_reply"):
return []
args = []
# Bare trigger. dm.py's parse_tags splits each --tag on '='; an empty
# value means "present". If the deployed dm.py requires a non-empty
# value for this key, use 'reply:expected=true' instead.
args += ["--tag", "reply:expected="]
if "timeout" in followup:
s = parse_followup_duration(followup["timeout"])
args += ["--tag", "reply:timeout=%d" % s]
if "nudges" in followup:
n = followup["nudges"]
if not isinstance(n, int) or isinstance(n, bool) or not 0 <= n <= 10:
raise ValueError("nudges must be int 0..10")
args += ["--tag", "reply:nudges=%d" % n]
if "escalate" in followup:
e = followup["escalate"]
if not isinstance(e, str) or not re.fullmatch(r"[a-z0-9_-]{1,64}", e):
raise ValueError("escalate must be an identity string")
args += ["--tag", "reply:escalate=%s" % e]
if "route" in followup:
r = followup["route"]
if not isinstance(r, str) or not re.fullmatch(r"[a-z0-9_-]{1,64}", r):
raise ValueError("route must be a route_id string")
args += ["--tag", "route:%s" % r]
# 'thread' is intentionally not settable from job JSON; it names a
# specific existing thread and is filled by the dispatcher when known.
return args
def send_dm(agent, target, message, dry_run=False, followup_tags=None,
allow_main_chat=False):
"""Send DM via dm.py. followup_tags: flat ['--tag', 'k=v', ...] list
from build_followup_tags(), or None. allow_main_chat passes the explicit
main-chat opt-in through to dm.py (sidechat-first policy)."""
if dry_run:
print(f"[DRY RUN] Would send to {agent} ({target}):")
if followup_tags:
print(f"[DRY RUN] With follow-up tags: {' '.join(followup_tags)}")
print(message[:200] + "..." if len(message) > 200 else message)
return "dry-run-id"
# Rate limit
if HAS_RATE_LIMITER:
rate_limit_wait(agent)
# Cryptographic attestation: only sign if explicitly requested by job config
# to avoid blowing up agent context windows with massive base64 SSH signature blocks.
signed_payload = None
if os.environ.get("JOB_REQUIRE_SIGNATURE") == "1":
dm_sign_sh = NETVM_ROOT / "bin" / "dm-sign.sh"
priv_key = Path(os.path.expanduser("~/.ssh/id_ed25519"))
if dm_sign_sh.exists() and priv_key.exists():
try:
sign_res = subprocess.run(
[str(dm_sign_sh), "--from", "super", "--key", str(priv_key), message],
capture_output=True, text=True, timeout=10
)
if sign_res.returncode == 0 and "-----BEGIN SSH SIGNATURE-----" in sign_res.stdout:
signed_payload = sign_res.stdout.strip()
id_m = re.search(r"\[id:([a-f0-9]+)\]", signed_payload)
proof_id = id_m.group(1) if id_m else None
if proof_id:
proof_data = {
"id": proof_id,
"signer": "super",
"target_agent": agent,
"target_conversation": target,
"raw_payload": signed_payload,
"ts": datetime.now(timezone.utc).isoformat()
}
try:
import urllib.request
req = urllib.request.Request(
"https://crypt.muse-dev.online/proofs",
data=json.dumps(proof_data).encode("utf-8"),
headers={"Content-Type": "application/json", "User-Agent": "job-dispatch/1.0"},
method="POST"
)
with urllib.request.urlopen(req, timeout=3) as resp:
pass
except Exception as pe:
sys.stderr.write(f"warning: proof registration to crypt.muse-dev.online failed: {pe}\n")
except Exception as se:
sys.stderr.write(f"warning: dm signing failed: {se}\n")
payload_to_send = signed_payload or message
# Try fast hybrid gateway send if target resolves to UUID
target_uuid = None
try:
import dm
target_uuid = dm.resolve_sidechat_target(target, agent)
except Exception:
pass
if target_uuid and re.fullmatch(r"[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}", target_uuid.lower()):
try:
import muse_hybrid
send_node = agent if agent in ("muse", "pip", "646", "opm") else "opm"
res, err = muse_hybrid.send_message(send_node, payload_to_send, thread_id=target_uuid, wait=0)
if res and not err:
m_id = None
id_m = re.search(r"\[id:([a-f0-9]+)\]", payload_to_send)
if id_m:
m_id = id_m.group(1)
log_entry = {
"type": "sent",
"id": m_id or (res.get("reply", {}).get("message_id") if isinstance(res, dict) else "gateway"),
"agent": "opm",
"to": agent,
"target": target,
"thread_uuid": target_uuid,
"transport": "gateway",
"verified": True,
"ts": datetime.now(timezone.utc).isoformat()
}
dm_log_path = NETVM_ROOT / "dm-log.jsonl"
with open(dm_log_path, "a", encoding="utf-8") as lf:
lf.write(json.dumps(log_entry) + "\n")
return m_id or "gateway-verified"
sys.stderr.write(f"warning: fast gateway send fallback: {err or res}\n")
except Exception as ge:
sys.stderr.write(f"warning: fast gateway send fallback: {ge}\n")
if signed_payload:
cmd = ([str(DM_PY), "send", "--agent", "opm", "--to", agent,
"--target", target, "--raw"]
+ (["--allow-main-chat"] if (target == "main" and allow_main_chat) else [])
+ (followup_tags or []) + [signed_payload])
else:
cmd = ([str(DM_PY), "send", "--agent", "opm", "--to", agent,
"--target", target]
+ (["--allow-main-chat"] if (target == "main" and allow_main_chat) else [])
+ (followup_tags or []) + [message])
result = subprocess.run(cmd, capture_output=True, text=True, timeout=120)
if result.returncode != 0:
print(f"DM send failed: {result.stderr}", file=sys.stderr)
return None
# Extract message ID from output (format: DM <id> ... SENT and VERIFIED)
output = result.stdout.strip()
m = re.search(r"DM\s+([a-f0-9]{8})", output)
return m.group(1) if m else "verified"
def create_sidechat(sender_agent, dry_run=False):
"""Create a sidechat via muse-chat-api.py in sender's context.
Returns True on success (browser now on new sidechat), False on failure."""
if dry_run:
print(f"[DRY RUN] Would create sidechat for {sender_agent}")
return True
cmd = [NETVM_EXEC, sender_agent, "--", "python3", str(CHAT_API),
"--account", sender_agent, "sidechat", "create"]
try:
result = subprocess.run(cmd, capture_output=True, text=True, timeout=90)
output = result.stdout.strip()
err = result.stderr.strip()
# Success if we see "Created:" (URL may be /thread/new placeholder)
if "Created:" in output:
print(f"Sidechat created", file=sys.stderr)
return True
print(f"Sidechat create failed. stdout: {output[:300]}", file=sys.stderr)
print(f"Sidechat create stderr: {err[:300]}", file=sys.stderr)
print(f"Return code: {result.returncode}", file=sys.stderr)
return False
except Exception as e:
print(f"Sidechat creation failed: {e}", file=sys.stderr)
return False
def send_to_current_chat(sender_agent, message, dry_run=False):
"""Send message to current chat via muse-chat-api.py (no navigation).
Used after sidechat create - browser is already on the new chat."""
if dry_run:
print(f"[DRY RUN] Would send to current chat: {message[:100]}...")
return "dry-run-id"
if HAS_RATE_LIMITER:
rate_limit_wait(sender_agent)
cmd = [NETVM_EXEC, sender_agent, "--", "python3", str(CHAT_API),
"--account", sender_agent, "send", message]
try:
result = subprocess.run(cmd, capture_output=True, text=True, timeout=60)
if result.returncode == 0:
return "sent-to-sidechat"
print(f"Send failed: {result.stderr[:200]}", file=sys.stderr)
return None
except Exception as e:
print(f"Send failed: {e}", file=sys.stderr)
return None
def main():
p = argparse.ArgumentParser(description="Dispatch a job by sending a DM to an agent.")
p.add_argument("job_name", help="Name of the job (without .json)")
p.add_argument("--dry-run", action="store_true", help="Print what would be sent without sending")
p.add_argument("--pipeline-run", default=os.environ.get("CHAIN_PIPELINE_RUN_ID"),
help="Pipeline run ID if running as part of a pipeline")
p.add_argument("--step-n", type=int, default=int(os.environ.get("CHAIN_STEP_N", "1")),
help="Step sequence number in the pipeline")
p.add_argument("--allow-main-chat", action="store_true",
help="Explicit opt-in: allow this job to dispatch to main chat (refused by default per sidechat-first policy)")
args = p.parse_args()
job_name = args.job_name
dry_run = args.dry_run
pipeline_run_id = args.pipeline_run
step_n = args.step_n
# Load job
job = load_job(job_name)
# Generate job_id
job_id = f"{job_name}-{datetime.now(timezone.utc).strftime('%Y%m%d-%H%M%S')}-{uuid.uuid4().hex[:8]}"
# Variables for template
variables = {
"job_id": job_id,
"job_name": job_name,
"date": datetime.now(timezone.utc).strftime("%Y-%m-%d"),
"datetime": datetime.now(timezone.utc).isoformat(),
"prev_job_id": os.environ.get("CHAIN_PREV_JOB_ID", ""),
"prev_result": os.environ.get("CHAIN_PREV_RESULT", ""),
"pipeline_run_id": pipeline_run_id or "",
"step_n": step_n,
}
# Follow-up tracking (opt-in via job JSON "followup" block; see helpers).
# The heartbeat job is a loopback health check and must never be tracked.
followup_cfg = job.get("followup")
followup_tags = []
if followup_cfg:
if job_name == HEARTBEAT_JOB_NAME:
print(f"Warning: job '{job_name}' must not use follow-up "
f"tracking (loopback); ignoring followup block",
file=sys.stderr)
log_event("job_followup_skipped",
{"job_id": job_id, "reason": "heartbeat_loopback"})
else:
try:
followup_tags = build_followup_tags(followup_cfg)
if followup_tags:
log_event("job_followup_armed",
{"job_id": job_id, "tags": followup_tags})
except ValueError as e:
print(f"Warning: invalid followup block: {e}; "
f"sending untagged", file=sys.stderr)
log_event("job_followup_invalid",
{"job_id": job_id, "error": str(e)})
# Render prompt
prompt_template = job.get("prompt_template", "")
if not prompt_template:
print(f"Error: Job '{job_name}' has no prompt_template", file=sys.stderr)
sys.exit(1)
rendered = render_prompt(prompt_template, variables)
# Determine recipient agent and target
agent = job.get("agent", "muse")
target = "main"
if pipeline_run_id:
if HAS_PIPELINE:
p_entry = pipeline_engine.get_pipeline(pipeline_run_id)
if not p_entry:
p_entry = pipeline_engine.create_pipeline(job_name, run_id=pipeline_run_id)
target = p_entry.get("target") or f"pipe-{pipeline_run_id.split('-')[-1]}"
else:
target = f"pipe-{pipeline_run_id.split('-')[-1]}"
elif job.get("sidechat", {}).get("create"):
sc_cfg = job.get("sidechat", {})
sc_name = render_prompt(sc_cfg.get("name_template", "job-{job_name}-{date}"), variables)
reuse_key = sc_cfg.get("reuse_key")
# Check if reuse_key exists in job-sidechats.json and thread is still alive
sc_state = load_sidechat_state()
reused_uuid = None
if reuse_key and reuse_key in sc_state:
val = sc_state[reuse_key]
cand_uuid = val.get("thread_uuid") if isinstance(val, dict) else val
if cand_uuid:
try:
import muse_hybrid
threads, err = muse_hybrid.get_threads(agent)
if not err and threads:
thread_ids = [t.get("session_id") for t in threads]
if cand_uuid in thread_ids:
reused_uuid = cand_uuid
except Exception:
pass
if reused_uuid:
target = reused_uuid
else:
# Spawn a brand new sidechat/channel via fast headless gateway!
channel_title = sc_name
try:
import muse_hybrid
res, err = muse_hybrid.start_session(agent, title=channel_title)
if res and not err and res.get("session_id"):
new_uuid = res.get("session_id")
key_to_save = reuse_key or sc_name
is_persistent = bool(reuse_key)
sc_state[key_to_save] = {
"thread_uuid": new_uuid,
"agent": agent,
"title": channel_title,
"type": "persistent" if is_persistent else "ephemeral",
"created_at": datetime.now(timezone.utc).isoformat()
}
save_sidechat_state(sc_state)
target = new_uuid
print(f"Spawned new sidechat channel '{channel_title}' ({new_uuid}) for {agent}")
else:
target = sc_name
except Exception as e:
sys.stderr.write(f"warning: fast gateway session-start exception ({e}), falling back to name {sc_name}\n")
target = sc_name
elif job.get("dm_target"):
target = job.get("dm_target").strip()
elif job.get("target"):
target = job.get("target").strip()
# Pre-dispatch approval & input check (auto-approve trusted; warn if blocked)
if not dry_run:
try:
import approvals
app_info = approvals.inspect_node_approvals(agent)
if app_info.get("has_pending"):
if app_info.get("is_trusted") and app_info.get("status") != "KEY_APPROVAL":
print(f"Pre-dispatch: auto-approving trusted request for {agent} ({app_info.get('target')})")
approvals.allow_node_approval(agent, always=True, caller="job-dispatch")
else:
print(f"Warning: Agent '{agent}' has untrusted pending approval ({app_info.get('target')}). Dispatch may stall.", file=sys.stderr)
log_event("job_dispatch_approval_blocked", {"job_id": job_id, "agent": agent, "target": app_info.get("target")})
elif app_info.get("status") == "INPUT_WAIT":
waits = app_info.get("input_waits", [])
w_desc = "; ".join(w.get("task", "") for w in waits)[:80]
print(f"Notice: Agent '{agent}' has task waiting for input ({w_desc}).", file=sys.stderr)
log_event("job_dispatch_agent_input_wait", {"job_id": job_id, "agent": agent, "waits": w_desc})
except Exception:
pass
# Sidechat-first policy (2026-10-04): refuse to dispatch to main chat
# unless the job explicitly opts in. Never fall back to main silently.
allow_main = bool(job.get("allow_main_chat")) or args.allow_main_chat
if target == "main" and not allow_main:
print(f"ERROR: Job '{job_name}' resolves to main chat; refusing by sidechat-first policy. "
f"Set a sidechat target (dm_target/target/sidechat.create) in the job JSON, "
f"set \"allow_main_chat\": true, or pass --allow-main-chat.", file=sys.stderr)
log_event("job_failed", {
"job_id": job_id,
"error": "main_chat_blocked_by_policy",
"pipeline_run_id": pipeline_run_id,
})
sys.exit(1)
# Log job_sent
log_event("job_sent", {
"job_id": job_id,
"job_name": job_name,
"agent": agent,
"target": target,
"dry_run": dry_run,
"pipeline_run_id": pipeline_run_id,
"step_n": step_n,
})
# Work-first envelope: executable swarm.spawn/followup.create at TOP and BOTTOM
# (see bin/prompt_envelope.py). Skipped when the job sets "skip_envelope": true
# (agents whose runtime lacks the enveloped tools, e.g. pip).
if not job.get("skip_envelope"):
import prompt_envelope
rendered = prompt_envelope.wrap(job_name, job_id, agent, target, rendered)
# Format as JOB DM
dm_message = f"[JOB {job_id}] {rendered}"
# Dispatch via dm.py (handles main or sidechat with auto-provisioning and verification)
msg_id = send_dm(agent, target, dm_message,
dry_run=dry_run, followup_tags=followup_tags,
allow_main_chat=allow_main)
if msg_id and not dry_run:
print(f"Dispatched job {job_id} to {agent}/{target} (DM: {msg_id})")
log_event("job_dispatched", {
"job_id": job_id,
"dm_id": msg_id,
"pipeline_run_id": pipeline_run_id,
"step_n": step_n,
})
if pipeline_run_id and HAS_PIPELINE:
pipeline_engine.record_step_dispatch(
pipeline_run_id, step_n, job_name, job_id, agent, target, dm_id=msg_id
)
sc_state = load_sidechat_state()
if target in sc_state:
val = sc_state[target]
t_uuid = val.get("thread_uuid") if isinstance(val, dict) else val
if t_uuid:
pipeline_engine.update_pipeline_thread(pipeline_run_id, t_uuid)
elif dry_run:
print(f"[DRY RUN] Job {job_id} would be dispatched to {agent}/{target}")
else:
print(f"Failed to dispatch job {job_id}", file=sys.stderr)
log_event("job_failed", {
"job_id": job_id,
"error": "dm_send_failed",
"pipeline_run_id": pipeline_run_id,
})
if pipeline_run_id and HAS_PIPELINE:
pipeline_engine.fail_pipeline(pipeline_run_id, "dm_send_failed")
sys.exit(1)
if __name__ == "__main__":
main()
-211
View File
@@ -1,211 +0,0 @@
#!/usr/bin/env python3
"""job-scheduler.py — NetVM unified job scheduler.
Evaluates cron schedules in jobs/*.json and triggers due jobs via job-dispatch.py.
Prevents duplicate dispatches using state watermarks in /home/super/Projects/NetVM/job-scheduler-state.json.
Usage:
python3 bin/job-scheduler.py run [--dry-run]
python3 bin/job-scheduler.py status
"""
import os
import sys
import glob
import json
import fcntl
import argparse
import subprocess
from datetime import datetime, timezone, timedelta
BASE = "/home/super/Projects/NetVM"
JOBS_DIR = os.path.join(BASE, "jobs")
BIN = os.path.join(BASE, "bin")
JOB_DISPATCH = os.path.join(BIN, "job-dispatch.py")
STATE_FILE = os.path.join(BASE, "job-scheduler-state.json")
LOCK_FILE = os.path.join(BASE, "job-scheduler.lock")
def utcnow():
return datetime.now(timezone.utc)
def parse_field(pattern, val):
if pattern == "*":
return True
for part in pattern.split(","):
if "/" in part:
sub = part.split("/")
step = int(sub[1])
base = sub[0]
start = 0 if base == "*" else int(base.split("-")[0])
end = 59 if base == "*" else int(base.split("-")[-1])
if start <= val <= end and (val - start) % step == 0:
return True
elif "-" in part:
s, e = map(int, part.split("-"))
if s <= val <= e:
return True
elif part.isdigit() and int(part) == val:
return True
return False
def cron_matches(expr, dt):
"""Check if 5-field cron expression matches datetime dt."""
parts = expr.strip().split()
if len(parts) != 5:
return False
m, h, dom, mon, dow = parts
dow_val = (dt.weekday() + 1) % 7 # 0=Sunday
return (parse_field(m, dt.minute) and
parse_field(h, dt.hour) and
parse_field(dom, dt.day) and
parse_field(mon, dt.month) and
(parse_field(dow, dt.weekday() + 1) or parse_field(dow, dow_val)))
def is_job_due(expr, now_dt, last_fired_dt=None):
"""Determine if a cron job is due within the recent 5-minute sampling window."""
if not expr or expr.strip().lower() == "manual":
return False
# Check minutes in the window [now - 4min, now]
matched_dt = None
for offset in range(5):
sample_dt = now_dt - timedelta(minutes=offset)
if cron_matches(expr, sample_dt):
matched_dt = sample_dt.replace(second=0, microsecond=0)
break
if not matched_dt:
return False
if last_fired_dt:
# If fired within 4 minutes of the matched slot, skip duplicate
diff_seconds = (now_dt - last_fired_dt).total_seconds()
# For hourly or longer jobs, prevent re-fire within 45 minutes
if " " in expr and expr.split()[0] != "*":
if diff_seconds < 2700:
return False
elif diff_seconds < 240:
return False
return True
def load_state():
try:
with open(STATE_FILE, "r") as f:
return json.load(f)
except Exception:
return {"jobs": {}, "last_run": None}
def save_state(state):
tmp = STATE_FILE + ".tmp"
with open(tmp, "w") as f:
json.dump(state, f, indent=2)
os.replace(tmp, STATE_FILE)
def do_run(dry_run=False):
now = utcnow()
now_iso = now.strftime("%Y-%m-%dT%H:%M:%SZ")
state = load_state()
jobs_state = state.setdefault("jobs", {})
job_files = sorted(glob.glob(os.path.join(JOBS_DIR, "*.json")))
dispatched = []
skipped = []
for jpath in job_files:
try:
with open(jpath, "r", encoding="utf-8") as f:
data = json.load(f)
except Exception:
continue
job_name = data.get("name") or os.path.basename(jpath).replace(".json", "")
schedule = data.get("schedule")
if not schedule or schedule.strip().lower() == "manual":
continue
j_st = jobs_state.get(job_name, {})
last_fired_str = j_st.get("last_fired")
last_fired_dt = None
if last_fired_str:
try:
last_fired_dt = datetime.fromisoformat(last_fired_str.replace("Z", "+00:00"))
except Exception:
pass
if is_job_due(schedule, now, last_fired_dt):
if dry_run:
print(f"[DRY RUN] Due job: {job_name} ({schedule})")
dispatched.append(job_name)
continue
cmd = [sys.executable, JOB_DISPATCH, job_name]
try:
r = subprocess.run(cmd, capture_output=True, text=True, timeout=180)
if r.returncode == 0:
dispatched.append(job_name)
jobs_state[job_name] = {
"last_fired": now_iso,
"schedule": schedule,
"status": "dispatched"
}
print(f"Dispatched job: {job_name} ({schedule})", file=sys.stderr)
else:
err = (r.stderr or r.stdout).strip()[-200:]
print(f"Failed to dispatch {job_name}: {err}", file=sys.stderr)
except Exception as e:
print(f"Exception dispatching {job_name}: {e}", file=sys.stderr)
else:
skipped.append(job_name)
if not dry_run:
state["last_run"] = now_iso
save_state(state)
result = {
"ok": True,
"dispatched": dispatched,
"dispatched_count": len(dispatched),
"evaluated_at": now_iso
}
return result
def main():
p = argparse.ArgumentParser(description="NetVM Unified Job Scheduler")
sub = p.add_subparsers(dest="cmd")
p_run = sub.add_parser("run", help="Evaluate schedules and dispatch due jobs")
p_run.add_argument("--dry-run", action="store_true", help="Print due jobs without dispatching")
sub.add_parser("status", help="Show scheduler state and last run")
args = p.parse_args()
cmd = args.cmd or "run"
if cmd == "status":
print(json.dumps(load_state(), indent=2))
return 0
if cmd == "run":
try:
lockfh = open(LOCK_FILE, "w")
fcntl.flock(lockfh, fcntl.LOCK_EX | fcntl.LOCK_NB)
except (OSError, IOError):
print(json.dumps({"ok": False, "skipped": "already running"}))
return 0
res = do_run(dry_run=getattr(args, "dry_run", False))
print(json.dumps(res))
return 0
return 0
if __name__ == "__main__":
sys.exit(main())
-94
View File
@@ -1,94 +0,0 @@
#!/usr/bin/env python3
"""Thread keepalive timer entry point (runs on bl).
Reads keepalive-config.json, ticks every registered thread:
ensure reachable -> ping if idle past threshold -> recreate if dead.
Designed to run from a systemd timer (e.g. every 5 minutes). Each tick is
idempotent and cheap: threads that are alive and recently active are a
single navigate+verify (no ping sent).
Usage:
keepalive-timer.py [--config PATH] [--state PATH] [--dry-run] [--status]
--status print the registry with idle times and exit (no changes)
Exit codes: 0 = all ok, 1 = one or more threads failed, 2 = config error.
"""
import argparse
import json
import os
import sys
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
from keepalive import Keepalive
NETVM_ROOT = os.environ.get("NETVM_ROOT", "/home/super/Projects/NetVM")
DEFAULT_CONFIG = os.path.join(NETVM_ROOT, "keepalive-config.json")
DEFAULT_STATE = os.path.join(NETVM_ROOT, "keepalive-threads.json")
def load_config(path):
try:
with open(path) as f:
cfg = json.load(f)
except FileNotFoundError:
print(f"keepalive: config not found: {path}", file=sys.stderr)
return None
except Exception as e:
print(f"keepalive: bad config {path}: {e}", file=sys.stderr)
return None
threads = cfg.get("threads", [])
if not isinstance(threads, list):
print("keepalive: config 'threads' must be a list", file=sys.stderr)
return None
return threads
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--config", default=DEFAULT_CONFIG)
ap.add_argument("--state", default=DEFAULT_STATE)
ap.add_argument("--dry-run", action="store_true")
ap.add_argument("--status", action="store_true")
args = ap.parse_args()
ka = Keepalive(state_path=args.state, dry_run=args.dry_run)
if args.status:
print(json.dumps(ka.status(), indent=2))
return 0
threads = load_config(args.config)
if threads is None:
return 2
results = {}
for t in threads:
key = t.get("key")
agent = t.get("agent", "opm")
if not key or not t.get("enabled", True):
continue
try:
status = ka.tick(
key,
agent,
idle_threshold_s=int(t.get("idle_threshold_s", 3600)),
max_retries=int(t.get("max_retries", 3)),
ping_message=t.get("ping_message"),
)
except Exception as e:
status = f"error: {e}"
ka.log_event("keepalive_tick_error",
{"key": key, "agent": agent, "error": str(e)[:200]})
results[key] = status
print(f"keepalive: {key}@{agent} -> {status}")
failed = [k for k, v in results.items()
if v in ("failed",) or v.startswith("error")]
return 1 if failed else 0
if __name__ == "__main__":
sys.exit(main())
-345
View File
@@ -1,345 +0,0 @@
#!/usr/bin/env python3
"""Generic thread keepalive for bl side chats.
Pattern proven by the heartbeat job overnight: a state file maps a stable
key -> thread UUID, and each tick navigates directly to
https://muse.ai/thread/<uuid> (UUID reuse fix in muse-chat-api.py
cmd_sidechat_use) instead of name-based sidebar lookup.
This module generalizes that pattern:
- ANY side chat can register for keepalive (not just job sidechats).
- Each tick: ensure the thread is reachable; ping it if idle past threshold;
recreate it if the stored UUID is dead.
- State lives in keepalive-threads.json (atomic write via tmp+rename).
Usage:
from keepalive import Keepalive
ka = Keepalive(state_path="/home/super/Projects/NetVM/keepalive-threads.json")
ka.ensure("ops-watch", agent="opm") # reuse or create
ka.tick("ops-watch", agent="opm",
ping_message="[keepalive] ops-watch {ts}",
idle_threshold_s=3600) # ping if idle
The timer entry point (keepalive-timer.py) drives this from a JSON config.
"""
import json
import os
import re
import subprocess
import sys
import time
from datetime import datetime, timezone
# ---------------------------------------------------------------------------
# Paths (overridable for tests)
# ---------------------------------------------------------------------------
NETVM_ROOT = os.environ.get("NETVM_ROOT", "/home/super/Projects/NetVM")
NETVM_EXEC = os.path.join(NETVM_ROOT, "bin", "netvm-exec.sh")
CHAT_API = os.path.join(NETVM_ROOT, "bin", "muse-chat-api.py")
DEFAULT_STATE = os.path.join(NETVM_ROOT, "keepalive-threads.json")
DEFAULT_LOG = os.path.join(NETVM_ROOT, "keepalive-log.jsonl")
UUID_RE = re.compile(
r"/thread/([0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-"
r"[0-9a-f]{4}-[0-9a-f]{12})"
)
def _utcnow():
return datetime.now(timezone.utc).isoformat()
def _ts():
return time.time()
# ---------------------------------------------------------------------------
# Registry
# ---------------------------------------------------------------------------
class Keepalive:
"""Thread keepalive registry + ensure/ping/check operations."""
def __init__(self, state_path=DEFAULT_STATE, log_path=DEFAULT_LOG,
dry_run=False):
self.state_path = state_path
self.log_path = log_path
self.dry_run = dry_run
# -- state ---------------------------------------------------------
def load_state(self):
if os.path.exists(self.state_path):
try:
with open(self.state_path) as f:
data = json.load(f)
return data if isinstance(data, dict) else {}
except Exception:
return {}
return {}
def save_state(self, state):
if self.dry_run:
return
tmp = self.state_path + ".tmp"
with open(tmp, "w") as f:
json.dump(state, f, indent=2)
os.replace(tmp, self.state_path)
def log_event(self, event_type, data):
if self.dry_run:
return
entry = {"ts": _utcnow(), "type": event_type, **data}
try:
with open(self.log_path, "a") as f:
f.write(json.dumps(entry) + "\n")
except Exception:
pass
# -- browser ops (mirrors job-dispatch.py; UUID navigation, not names) --
def _run(self, agent, *api_args, timeout=60):
"""Run muse-chat-api.py in the agent's netns via netvm-exec.sh."""
cmd = [NETVM_EXEC, agent, "--", "python3", CHAT_API,
"--account", agent] + list(api_args)
try:
r = subprocess.run(cmd, capture_output=True, text=True,
timeout=timeout)
return r.returncode, r.stdout.strip(), r.stderr.strip()
except Exception as e:
return -1, "", str(e)
def navigate_to_uuid(self, agent, thread_uuid):
"""Go directly to https://muse.ai/thread/<uuid>.
This is the heartbeat UUID-reuse fix: cmd_sidechat_use navigates by
URL for UUID args instead of searching sidebar titles by name.
Returns True if the command succeeded.
"""
if self.dry_run:
print(f"[DRY] navigate {agent} -> {thread_uuid}")
return True
rc, out, err = self._run(agent, "sidechat", "use", thread_uuid,
timeout=45)
return rc == 0
def current_thread_uuid(self, agent):
"""Read the browser's current URL and extract the thread UUID."""
if self.dry_run:
return None
rc, out, err = self._run(agent, "url", timeout=30)
if rc != 0:
return None
m = UUID_RE.search(out or "")
return m.group(1) if m else None
def create_sidechat(self, agent, name=None):
"""Create a new sidechat in the agent's account. Returns True."""
if self.dry_run:
print(f"[DRY] create sidechat for {agent}")
return True
rc, out, err = self._run(agent, "sidechat", "create", timeout=90)
return rc == 0 and "Created:" in (out or "")
def send_message(self, agent, message):
"""Send a message to the current chat. Returns True on success."""
if self.dry_run:
print(f"[DRY] send to {agent}: {message[:80]}")
return True
rc, out, err = self._run(agent, "send", message, timeout=60)
return rc == 0
# -- keepalive ops ---------------------------------------------------
def check(self, key, agent):
"""Verify the registered thread is reachable.
Navigates to the stored UUID and confirms the browser lands on it.
Returns (ok, thread_uuid). Updates last_verified_ts on success.
"""
state = self.load_state()
rec = state.get(key)
if not rec or not rec.get("thread_uuid"):
return False, None
uuid = rec["thread_uuid"]
if not self.navigate_to_uuid(agent, uuid):
self.log_event("keepalive_check_failed",
{"key": key, "agent": agent, "thread_uuid": uuid,
"reason": "navigate_failed"})
return False, uuid
cur = self.current_thread_uuid(agent)
if cur == uuid:
rec["last_verified_ts"] = _ts()
rec["consecutive_failures"] = 0
state[key] = rec
self.save_state(state)
self.log_event("keepalive_check_ok",
{"key": key, "agent": agent, "thread_uuid": uuid})
return True, uuid
self.log_event("keepalive_check_failed",
{"key": key, "agent": agent, "thread_uuid": uuid,
"reason": "url_mismatch", "current": cur})
return False, uuid
def ensure(self, key, agent, name=None):
"""Ensure the thread exists and is reachable; create if needed.
Returns (ok, thread_uuid). Mirrors the job-dispatch.py reuse-or-
create flow: try stored UUID first, fall back to creation, then
capture the new UUID from the browser URL.
"""
state = self.load_state()
rec = state.get(key)
if rec and rec.get("thread_uuid"):
ok, uuid = self.check(key, agent)
if ok:
return True, uuid
# Stored thread is dead — fall through to recreate.
print(f"keepalive: stored thread for {key} unreachable, "
f"recreating", file=sys.stderr)
# Create a fresh sidechat.
if not self.create_sidechat(agent, name=name):
self.log_event("keepalive_create_failed",
{"key": key, "agent": agent})
return False, None
# Capture the new thread UUID from the browser URL.
uuid = None
if not self.dry_run:
for _ in range(15):
time.sleep(1)
uuid = self.current_thread_uuid(agent)
if uuid:
break
else:
uuid = "dry-run-uuid"
if not uuid:
self.log_event("keepalive_create_failed",
{"key": key, "agent": agent,
"reason": "uuid_capture_failed"})
return False, None
now = _ts()
state[key] = {
"thread_uuid": uuid,
"agent": agent,
"created_ts": now,
"last_ping_ts": 0,
"last_verified_ts": now,
"last_activity_ts": now,
"ping_count": 0,
"consecutive_failures": 0,
}
self.save_state(state)
self.log_event("keepalive_created",
{"key": key, "agent": agent, "thread_uuid": uuid})
return True, uuid
def ping(self, key, agent, message=None):
"""Send a keepalive ping into the thread.
Navigates to the thread first (cheap no-op if already there),
then sends the ping message. Updates last_ping_ts / ping_count.
"""
state = self.load_state()
rec = state.get(key)
if not rec or not rec.get("thread_uuid"):
return False
uuid = rec["thread_uuid"]
if not self.navigate_to_uuid(agent, uuid):
return False
msg = (message or "[keepalive:{key}] tick {ts} {uuid}").format(
key=key, ts=_utcnow(), uuid=uuid)
if not self.send_message(agent, msg):
self.log_event("keepalive_ping_failed",
{"key": key, "agent": agent, "thread_uuid": uuid})
return False
now = _ts()
rec["last_ping_ts"] = now
rec["last_activity_ts"] = now
rec["ping_count"] = rec.get("ping_count", 0) + 1
state[key] = rec
self.save_state(state)
self.log_event("keepalive_ping",
{"key": key, "agent": agent, "thread_uuid": uuid,
"ping_count": rec["ping_count"]})
return True
def note_activity(self, key):
"""Record external activity (e.g. a job just sent to the thread)
so the idle timer doesn't ping unnecessarily."""
state = self.load_state()
rec = state.get(key)
if rec:
rec["last_activity_ts"] = _ts()
state[key] = rec
self.save_state(state)
def tick(self, key, agent, idle_threshold_s=3600, max_retries=3,
ping_message=None):
"""One keepalive tick for a registered thread.
- Ensures the thread exists (recreate if dead).
- Pings only if idle longer than idle_threshold_s.
- After max_retries consecutive failures, forces recreation.
Returns a status string: ok | pinged | recreated | failed.
"""
state = self.load_state()
rec = state.get(key)
if not rec or not rec.get("thread_uuid"):
ok, _ = self.ensure(key, agent)
return "recreated" if ok else "failed"
failures = rec.get("consecutive_failures", 0)
if failures >= max_retries:
# Force recreation: drop the dead UUID and re-ensure.
self.log_event("keepalive_force_recreate",
{"key": key, "agent": agent,
"failures": failures,
"old_uuid": rec.get("thread_uuid")})
rec["thread_uuid"] = None
state[key] = rec
self.save_state(state)
ok, _ = self.ensure(key, agent)
return "recreated" if ok else "failed"
ok, _ = self.check(key, agent)
if not ok:
state = self.load_state()
rec = state.get(key, {})
rec["consecutive_failures"] = failures + 1
state[key] = rec
self.save_state(state)
return "failed"
now = _ts()
idle_for = now - rec.get("last_activity_ts", 0)
if idle_for >= idle_threshold_s:
if self.ping(key, agent, message=ping_message):
return "pinged"
return "failed"
return "ok"
def unregister(self, key):
"""Remove a thread from keepalive (does not delete the sidechat)."""
state = self.load_state()
if key in state:
del state[key]
self.save_state(state)
self.log_event("keepalive_unregistered", {"key": key})
return True
return False
def status(self):
"""Return the full registry with computed idle times."""
state = self.load_state()
now = _ts()
out = {}
for key, rec in state.items():
r = dict(rec)
r["idle_s"] = int(now - rec.get("last_activity_ts", now))
out[key] = r
return out
-707
View File
@@ -1,707 +0,0 @@
#!/usr/bin/env python3
"""kpi.py — NetVM Fleet KPI, Spend Monitor & Runtime Preservation Engine.
Monitors:
- Calls / DMs dispatched and verified (from dm-log.jsonl)
- Token quota spend & remaining (weekly limit % and extra tokens)
- Active subagent sessions and tmux muse workers
- Uptime vs actual problems fixed (Efficiency Index)
- Route health (WARP wireguard, CDP, tmux sockets)
- Runtime preservation advisories (guiding agents to offload work to tmux/subagents)
"""
from __future__ import annotations
import argparse
import json
import os
import re
import subprocess
import sys
import time
from dataclasses import asdict, dataclass
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict, List, Optional
REPO_ROOT = Path(__file__).resolve().parent.parent
BIN_DIR = REPO_ROOT / "bin"
DM_LOG = REPO_ROOT / "dm-log.jsonl"
JOBS_DIR = REPO_ROOT / "jobs"
SUBAGENTS_FILE = REPO_ROOT / "subagent-sessions.json"
VALID_NODES = ["muse", "pip", "646", "opm", "dev", "def"]
@dataclass
class AgentKPI:
node: str
weekly_used_pct: Optional[int]
extra_tokens_remaining: str
is_blocked: bool
calls_sent: int
calls_verified: int
jobs_assigned: int
jobs_completed: int
subagents_active: int
tmux_workers_active: int
uptime_hours: float
route_status: str
efficiency_index: float
efficiency_rating: str
preservation_advisory: str
def to_dict(self) -> Dict[str, Any]:
return asdict(self)
def get_agent_dm_metrics(node: str, window_hours: Optional[float] = None) -> Dict[str, int]:
"""Calculate outbound messages, sends, and verified deliveries from dm-log.jsonl."""
if not DM_LOG.exists():
return {"sent": 0, "verified": 0, "total_events": 0}
cutoff = None
if window_hours:
cutoff = datetime.now(timezone.utc).timestamp() - (window_hours * 3600)
sent_ids = set()
verified_ids = set()
total_events = 0
try:
with open(DM_LOG, "r", encoding="utf-8") as f:
for line in f:
line = line.strip()
if not line:
continue
try:
entry = json.loads(line)
except Exception:
continue
if entry.get("agent") != node:
continue
if cutoff:
ts = entry.get("ts")
if ts:
try:
dt = datetime.fromisoformat(ts.replace("Z", "+00:00"))
if dt.timestamp() < cutoff:
continue
except Exception:
pass
total_events += 1
mid = entry.get("id")
etype = entry.get("type")
if etype in ("send_start", "send_done"):
if mid:
sent_ids.add(mid)
elif etype == "verified":
if mid:
verified_ids.add(mid)
except Exception:
pass
return {
"sent": len(sent_ids),
"verified": len(verified_ids),
"total_events": total_events,
}
def get_agent_job_metrics(node: str) -> Dict[str, int]:
"""Calculate total jobs assigned and completed for an agent."""
assigned = 0
completed = 0
if not JOBS_DIR.exists():
return {"assigned": 0, "completed": 0}
try:
for p in JOBS_DIR.glob("*.json"):
try:
with open(p, "r", encoding="utf-8") as f:
data = json.load(f)
if data.get("agent") == node:
assigned += 1
# If output or status has result
if data.get("status") == "completed" or data.get("result"):
completed += 1
except Exception:
continue
except Exception:
pass
return {"assigned": assigned, "completed": completed}
def get_agent_subagent_count(node: str) -> int:
"""Get active subagent sessions for a node from subagent-sessions.json."""
if not SUBAGENTS_FILE.exists():
return 0
try:
with open(SUBAGENTS_FILE, "r", encoding="utf-8") as f:
data = json.load(f)
if not isinstance(data, dict):
return 0
return sum(1 for s in data.values() if s.get("parent") == node and s.get("status") == "active")
except Exception:
return 0
def get_agent_tmux_workers(node: str) -> List[str]:
"""Get running tmux sessions for an agent (shared and netns socket)."""
sessions = []
# 1. Per-node socket
sock = f"/tmp/tmux-{node}.sock"
if os.path.exists(sock):
try:
r = subprocess.run(["tmux", "-S", sock, "list-sessions", "-F", "#{session_name}"], capture_output=True, text=True, timeout=2)
if r.returncode == 0 and r.stdout.strip():
sessions.extend(line.strip() for line in r.stdout.splitlines() if line.strip())
except Exception:
pass
# 2. Shared socket filtering sessions containing node name
shared_sock = "/tmp/tmux-muse.sock"
if os.path.exists(shared_sock):
try:
r = subprocess.run(["tmux", "-S", shared_sock, "list-sessions", "-F", "#{session_name}"], capture_output=True, text=True, timeout=2)
if r.returncode == 0 and r.stdout.strip():
for s in r.stdout.splitlines():
s = s.strip()
if s and (node in s or s.startswith(f"{node}-") or s == "swarm-worker"):
if s not in sessions:
sessions.append(s)
except Exception:
pass
return sessions
def get_agent_uptime_hours(node: str) -> float:
"""Calculate browser process uptime in hours."""
try:
# Search for chromium process matching user-data-dir or node
cmd = ["pgrep", "-f", f"chrome-box launch {node}"]
r = subprocess.run(cmd, capture_output=True, text=True, timeout=2)
pids = r.stdout.strip().split()
if not pids:
cmd = ["pgrep", "-f", f"profiles/{node}"]
r = subprocess.run(cmd, capture_output=True, text=True, timeout=2)
pids = r.stdout.strip().split()
if pids:
pid = pids[0]
# Read /proc/<pid>/stat starttime
stat_path = Path(f"/proc/{pid}/stat")
if stat_path.exists():
stat_content = stat_path.read_text().split()
# field 22 is starttime (in clock ticks after boot)
start_ticks = int(stat_content[21])
clk_tck = os.sysconf(os.sysconf_names["SC_CLK_TCK"])
with open("/proc/uptime", "r") as f:
uptime_sec = float(f.read().split()[0])
process_age_sec = uptime_sec - (start_ticks / clk_tck)
return round(max(0.0, process_age_sec / 3600.0), 2)
except Exception:
pass
return 0.0
def check_node_routes(node: str) -> str:
"""Check connectivity route for a node (netns + CDP)."""
# 1. Check netns
netns_path = Path(f"/var/run/netns/warp-{node}")
if not netns_path.exists():
return "NO_NETNS"
# 2. Check CDP page connection
try:
try:
from approvals import get_node_pages
except ImportError:
sys.path.insert(0, str(BIN_DIR))
from approvals import get_node_pages
pages = get_node_pages(node, timeout=2.0)
if pages:
return "ONLINE"
except Exception:
pass
return "DEGRADED"
def calculate_efficiency(
jobs_done: int,
subagents_active: int,
tmux_workers: int,
calls_verified: int,
weekly_used_pct: Optional[int],
uptime_hours: float,
) -> tuple[float, str]:
"""
Composite efficiency score.
Higher is better: measures actual work produced (jobs + subagents + workers + verified comms)
relative to quota burned and uptime elapsed.
"""
work_units = (jobs_done * 5.0) + (subagents_active * 3.0) + (tmux_workers * 4.0) + (calls_verified * 0.5)
burn_cost = max(1.0, (weekly_used_pct or 10) * 0.2)
# Base index
index = round(work_units / burn_cost, 2)
# Classify
if uptime_hours > 2.0 and subagents_active == 0 and tmux_workers == 0 and jobs_done == 0:
rating = "VANITY_IDLE"
elif index >= 3.0:
rating = "HIGH_EFFICIENCY"
elif index >= 1.0:
rating = "PRODUCTIVE"
elif index >= 0.4:
rating = "MODERATE"
else:
rating = "LOW_EFFICIENCY"
return index, rating
def generate_preservation_advisory(
node: str,
weekly_used_pct: Optional[int],
extra_tokens_remaining: str,
subagents_active: int,
tmux_workers: int,
rating: str,
) -> str:
"""Generate prescriptive runtime preservation instructions for the agent."""
tips = []
pct = weekly_used_pct or 0
if pct >= 95 or "0 tokens left" in extra_tokens_remaining:
return "CRITICAL: Quota exhausted. Do NOT send chat messages. Salvage via 'box onboard start <new_node> --for %s'." % node
if pct >= 70:
tips.append("Quota > 70%%: Cease prose chatter; offload tasks to background tmux workers.")
if subagents_active == 0 and tmux_workers == 0:
tips.append("Spawn subagents with isolated context ('box subagent spawn') or tmux muse workers.")
if rating in ("VANITY_IDLE", "LOW_EFFICIENCY"):
tips.append("Uptime without worker execution drains quota. Mandate: split goals into executable jobs.")
if not tips:
tips.append("Runtime healthy. Maintain worker-first execution strategy.")
return " ".join(tips)
def get_agent_kpi(node: str, usage_cache: Optional[Dict[str, Any]] = None) -> AgentKPI:
"""Collect comprehensive KPI metrics for a single NetVM node."""
# 1. Quota & Usage
usage = usage_cache
if usage is None:
try:
try:
import invite
except ImportError:
sys.path.insert(0, str(BIN_DIR))
import invite
u = invite.get_usage(node)
if isinstance(u, dict) and u.get("ok"):
usage = u
except Exception:
pass
if usage is None:
usage = {}
weekly_pct = usage.get("weekly_used_pct")
extra_left = usage.get("additional_left") or usage.get("extra_tokens_remaining") or "Unknown"
is_blocked = bool(usage.get("is_blocked")) or (weekly_pct is not None and weekly_pct >= 100 and "0 tokens left" in extra_left)
# 2. Activity metrics
dm_metrics = get_agent_dm_metrics(node)
job_metrics = get_agent_job_metrics(node)
subagents = get_agent_subagent_count(node)
tmux_sessions = get_agent_tmux_workers(node)
uptime = get_agent_uptime_hours(node)
route_status = check_node_routes(node)
# 3. Efficiency
eff_idx, eff_rating = calculate_efficiency(
jobs_done=job_metrics["completed"],
subagents_active=subagents,
tmux_workers=len(tmux_sessions),
calls_verified=dm_metrics["verified"],
weekly_used_pct=weekly_pct,
uptime_hours=uptime,
)
# 4. Advisory
advisory = generate_preservation_advisory(
node=node,
weekly_used_pct=weekly_pct,
extra_tokens_remaining=extra_left,
subagents_active=subagents,
tmux_workers=len(tmux_sessions),
rating=eff_rating,
)
return AgentKPI(
node=node,
weekly_used_pct=weekly_pct,
extra_tokens_remaining=extra_left,
is_blocked=is_blocked,
calls_sent=dm_metrics["sent"],
calls_verified=dm_metrics["verified"],
jobs_assigned=job_metrics["assigned"],
jobs_completed=job_metrics["completed"],
subagents_active=subagents,
tmux_workers_active=len(tmux_sessions),
uptime_hours=uptime,
route_status=route_status,
efficiency_index=eff_idx,
efficiency_rating=eff_rating,
preservation_advisory=advisory,
)
def fleet_kpi(nodes: Optional[List[str]] = None) -> Dict[str, AgentKPI]:
"""Collect KPI metrics across all fleet agents."""
target_nodes = nodes or VALID_NODES
# Fetch usage in bulk
usage_map = {}
try:
import invite
raw_usage = invite.fleet_usage(target_nodes)
if isinstance(raw_usage, dict):
usage_map = raw_usage
except Exception:
pass
results = {}
for n in target_nodes:
results[n] = get_agent_kpi(n, usage_cache=usage_map.get(n))
return results
def get_live_advisory_block(node: str) -> str:
"""Generate Markdown prompt envelope block ready for job injection."""
kpi = get_agent_kpi(node)
quota_str = f"{kpi.weekly_used_pct}% weekly limit used" if kpi.weekly_used_pct is not None else "quota active"
tokens_str = kpi.extra_tokens_remaining
lines = [
"---- BOX PERFORMANCE & RUNTIME ADVISORY ----",
f"AGENT: @{kpi.node} | QUOTA: {quota_str} ({tokens_str}) | UPTIME: {kpi.uptime_hours}h",
f"WORK UNITS: {kpi.jobs_completed} jobs finished | {kpi.subagents_active} subagents | {kpi.tmux_workers_active} tmux workers",
f"EFFICIENCY: {kpi.efficiency_rating} (Index: {kpi.efficiency_index}) | ROUTES: {kpi.route_status}",
f"RUNTIME MANDATE: {kpi.preservation_advisory}",
"Offload long operations to subagents or tmux muse workers to maximize problem-fixing per token.",
]
return "\n".join(lines)
NODE_SIDECHATS = {
"646": "646 tasks",
"opm": "heartbeat",
"pip": "646-pip-coord",
"dev": "dev-coord",
"def": "def-coord",
"muse": "646-muse-coord",
}
def find_pending_work_for_node(node: str) -> Optional[Dict[str, Any]]:
"""Find assigned pending job or swarm slot for an agent node."""
# 1. Look for node-specific auto-work jobs
if JOBS_DIR.exists():
candidates = sorted(list(JOBS_DIR.glob(f"auto-work-{node}-*.json")) + list(JOBS_DIR.glob(f"{node}-*.json")))
for c in candidates:
try:
with open(c, "r", encoding="utf-8") as f:
data = json.load(f)
job_agent = data.get("agent")
if job_agent and job_agent != node:
continue
job_name = c.stem
return {
"type": "job",
"name": job_name,
"path": str(c),
"cmd": f"{sys.executable} {BIN_DIR}/job-dispatch.py {job_name}",
}
except Exception:
continue
# 2. Check pending swarm slots
try:
from swarm_worker.poller import find_pending_slots
slots = find_pending_slots()
if slots:
slot = slots[0]
sw_id = slot.get("swarm_id", "swarm")
idx = slot.get("slot_index", 0)
return {
"type": "swarm",
"name": f"swarm-{sw_id}-s{idx}",
"path": None,
"cmd": f"{sys.executable} {BIN_DIR}/swarm_worker/daemon.py",
}
except Exception:
pass
return None
def auto_spawn_workers(nodes: Optional[List[str]] = None, dry_run: bool = False) -> List[Dict[str, Any]]:
"""Reconcile idle agents and auto-spawn background tmux workers to execute pending work."""
target_nodes = nodes or VALID_NODES
results = []
for node in target_nodes:
# Check active tmux workers for this node
active_tmux = len(get_agent_tmux_workers(node))
if active_tmux > 0:
results.append({
"node": node,
"action": "skip",
"reason": f"Active tmux worker already running ({active_tmux})",
})
continue
# Check work availability
work = find_pending_work_for_node(node)
if not work:
results.append({
"node": node,
"action": "idle",
"reason": "No pending jobs or swarm slots",
})
continue
session_label = f"worker-{work['name'][:18]}"
cmd_to_run = f"{work['cmd']} > /tmp/tmux-{node}-{session_label}.log 2>&1"
if dry_run:
results.append({
"node": node,
"action": "would_spawn",
"session": session_label,
"work_type": work["type"],
"work_name": work["name"],
"command": work["cmd"],
})
continue
# Execute spawn
spawn_res = spawn_tmux_worker(node, session_label, cmd_to_run)
if spawn_res.get("ok"):
# Send sidechat notification
try:
from invite_handler import send_loopback_notice
chat = NODE_SIDECHATS.get(node, "646 tasks")
msg = f"[BOX-AUTO-WORKER] Spawned background tmux worker '{session_label}' executing {work['type']} ({work['name']}). Logs at /tmp/tmux-{node}-{session_label}.log"
send_loopback_notice(recipient=node, target=chat, message=msg)
except Exception:
pass
results.append({
"node": node,
"action": "spawned",
"session": session_label,
"work_type": work["type"],
"work_name": work["name"],
"command": work["cmd"],
})
else:
results.append({
"node": node,
"action": "error",
"error": spawn_res.get("error", "Unknown spawn error"),
})
return results
def spawn_tmux_worker(node: str, session: str, command: str) -> Dict[str, Any]:
"""Spawn an autonomous tmux worker session on the agent's netns or shared socket."""
# Ensure session name is prefixed
clean_session = f"{node}-{session}" if not session.startswith(f"{node}-") else session
try:
from subagent_tracker import register_session
except ImportError:
sys.path.insert(0, str(BIN_DIR))
from subagent_tracker import register_session
muse_tmux = BIN_DIR / "muse-tmux.py"
if not muse_tmux.exists():
return {"ok": False, "error": "muse-tmux.py not found"}
# Execute via muse-tmux.py
cmd = [
sys.executable,
str(muse_tmux),
"new",
clean_session,
"--node",
node,
"--command",
command,
]
res = subprocess.run(cmd, capture_output=True, text=True, timeout=10)
if res.returncode != 0:
# Fallback to shared socket
cmd_shared = [
sys.executable,
str(muse_tmux),
"new",
clean_session,
"--command",
command,
]
res = subprocess.run(cmd_shared, capture_output=True, text=True, timeout=10)
if res.returncode != 0:
return {"ok": False, "error": res.stderr.strip() or res.stdout.strip()}
# Register in subagent tracker
sid = f"tmux-{clean_session}-{int(time.time())}"
register_session(parent=node, session_id=sid, title=f"tmux-worker-{clean_session}", prompt=command)
return {
"ok": True,
"node": node,
"session": clean_session,
"session_id": sid,
"command": command,
"message": f"Spawned tmux worker '{clean_session}' for @{node}. Running in background.",
}
def main():
parser = argparse.ArgumentParser(description="NetVM Fleet KPI, Spend Monitor & Runtime Preservation Engine")
subparsers = parser.add_subparsers(dest="command")
p_status = subparsers.add_parser("status", help="Show fleet KPI metrics table")
p_status.add_argument("--node", choices=VALID_NODES, help="Filter by node")
p_status.add_argument("--json", action="store_true", help="Emit JSON output")
p_report = subparsers.add_parser("report", help="Detailed KPI report for a specific node")
p_report.add_argument("node", choices=VALID_NODES, help="Target node")
p_report.add_argument("--json", action="store_true")
p_routes = subparsers.add_parser("routes", help="Verify network and CDP routes across nodes")
p_routes.add_argument("--json", action="store_true")
p_block = subparsers.add_parser("prompt-block", help="Generate live prompt envelope block for node")
p_block.add_argument("node", choices=VALID_NODES, help="Target node")
p_spawn = subparsers.add_parser("spawn-worker", help="Spawn autonomous background tmux worker session")
p_spawn.add_argument("node", choices=VALID_NODES, help="Agent node")
p_spawn.add_argument("session", help="Session label")
p_spawn.add_argument("worker_command", help="Command to execute inside worker")
p_autospawn = subparsers.add_parser("auto-spawn", help="Auto-spawn background tmux workers for idle nodes with pending work")
p_autospawn.add_argument("--node", choices=VALID_NODES, default=None, help="Filter by node")
p_autospawn.add_argument("--dry-run", action="store_true", help="Report what would be spawned without executing")
p_autospawn.add_argument("--json", action="store_true")
args = parser.parse_args()
if args.command in (None, "status"):
nodes = [args.node] if getattr(args, "node", None) else VALID_NODES
kpis = fleet_kpi(nodes)
if getattr(args, "json", False):
print(json.dumps({k: v.to_dict() for k, v in kpis.items()}, indent=2))
return
print("\n=== NETVM FLEET KPI & RUNTIME PRESERVATION DASHBOARD ===\n")
header = f"{'NODE':<6} {'QUOTA':<10} {'CALLS':<12} {'JOBS':<10} {'SUBAGENTS':<11} {'TMUX':<6} {'UPTIME':<8} {'ROUTE':<9} {'EFFICIENCY':<15}"
sep = f"{'────':<6} {'─────────':<10} {'───────────':<12} {'─────────':<10} {'──────────':<11} {'────':<6} {'──────':<8} {'───────':<9} {'──────────────':<15}"
print(header)
print(sep)
for n in nodes:
k = kpis.get(n)
if not k:
continue
q_str = f"{k.weekly_used_pct}%" if k.weekly_used_pct is not None else "Active"
c_str = f"{k.calls_sent} ({k.calls_verified}v)"
j_str = f"{k.jobs_completed}/{k.jobs_assigned}"
sub_str = str(k.subagents_active)
tmux_str = str(k.tmux_workers_active)
up_str = f"{k.uptime_hours}h"
print(f"{k.node:<6} {q_str:<10} {c_str:<12} {j_str:<10} {sub_str:<11} {tmux_str:<6} {up_str:<8} {k.route_status:<9} {k.efficiency_rating:<15}")
print("\nRun 'box kpi report <node>' for prescriptive runtime preservation advisories.\n")
elif args.command == "report":
kpi = get_agent_kpi(args.node)
if args.json:
print(json.dumps(kpi.to_dict(), indent=2))
return
print(f"\n=== KPI & RUNTIME REPORT: @{kpi.node.upper()} ===")
print(f" Weekly Quota: {kpi.weekly_used_pct}% used")
print(f" Extra Tokens: {kpi.extra_tokens_remaining}")
print(f" Blocked Status: {'YES (LIMIT REACHED)' if kpi.is_blocked else 'NO (HEALTHY)'}")
print(f" Messages / Calls: {kpi.calls_sent} sent ({kpi.calls_verified} verified delivered)")
print(f" Jobs Dispatched: {kpi.jobs_completed} completed / {kpi.jobs_assigned} assigned")
print(f" Active Subagents: {kpi.subagents_active}")
print(f" Active Tmux Workers:{kpi.tmux_workers_active}")
print(f" Process Uptime: {kpi.uptime_hours} hours")
print(f" Route Health: {kpi.route_status}")
print(f" Efficiency Index: {kpi.efficiency_index} ({kpi.efficiency_rating})")
print(f"\n [RUNTIME PRESERVATION ADVISORY]\n {kpi.preservation_advisory}\n")
elif args.command == "routes":
routes = {n: check_node_routes(n) for n in VALID_NODES}
if args.json:
print(json.dumps(routes, indent=2))
else:
print("\n=== NETVM ROUTE HEALTH ===")
for n, st in routes.items():
print(f" @{n:<6} : {st}")
print()
elif args.command == "prompt-block":
print(get_live_advisory_block(args.node))
elif args.command == "spawn-worker":
res = spawn_tmux_worker(args.node, args.session, args.worker_command)
if args.json:
print(json.dumps(res, indent=2))
else:
if res.get("ok"):
print(f"✔ {res.get('message')}")
else:
print(f"✘ Failed to spawn worker: {res.get('error')}", file=sys.stderr)
sys.exit(1)
elif args.command == "auto-spawn":
nodes = [args.node] if getattr(args, "node", None) else None
results = auto_spawn_workers(nodes=nodes, dry_run=args.dry_run)
if args.json:
print(json.dumps(results, indent=2))
return
print(f"\n=== AUTO-SPAWN WORKER RECONCILIATION {'(DRY-RUN)' if args.dry_run else ''} ===")
for r in results:
n = r.get("node")
act = r.get("action")
if act == "spawned":
print(f" ✔ @{n:<5} : SPAWNED session '{r.get('session')}' ({r.get('work_type')}: {r.get('work_name')})")
elif act == "would_spawn":
print(f" ? @{n:<5} : WOULD SPAWN session '{r.get('session')}' ({r.get('work_type')}: {r.get('work_name')})")
elif act == "skip":
print(f" - @{n:<5} : SKIP ({r.get('reason')})")
elif act == "idle":
print(f" - @{n:<5} : IDLE ({r.get('reason')})")
elif act == "error":
print(f" ✘ @{n:<5} : ERROR ({r.get('error')})")
print()
if __name__ == "__main__":
main()
-1
View File
@@ -1 +0,0 @@
/home/super/Projects/NetVM/bin/docs-lookup.py
-200
View File
@@ -1,200 +0,0 @@
#!/usr/bin/env python3
"""lookup_engine.py — Shared Zero-Downtime Hot-Reloading Pattern & Schema Engine.
Provides authoritative runtime access to lookup_internal/ databases for
daemons (response-harvester, self_main_loop, job-dispatch), CLI commands,
and agents.
Features:
- Dynamic mtime-based zero-downtime hot reloading of compiled regex patterns.
- Pre-flight soft validation of outbound agent sentences and work orders.
- Direct helper functions for core protocol regexes (RESULT, VERB, TOOL, etc.).
"""
import json
import os
import re
import sys
from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple
NETVM_ROOT = Path(__file__).resolve().parent.parent
LOOKUP_INTERNAL = NETVM_ROOT / "lookup_internal"
if not LOOKUP_INTERNAL.exists() and (NETVM_ROOT / "docs_internal").exists():
LOOKUP_INTERNAL = NETVM_ROOT / "docs_internal"
# In-memory cache structures with modification timestamps
_CACHE_MTIMES: Dict[str, float] = {}
_RAW_CACHE: Dict[str, Any] = {}
_COMPILED_PATTERNS: Dict[str, re.Pattern] = {}
# Fallback hardcoded regexes in case files are missing or unreadable
_FALLBACK_RESULT_RE = re.compile(r"\[RESULT\s+([A-Za-z0-9_/-]+)\]\s*(.*?)(?=\[RESULT\s|\Z)", re.S)
_FALLBACK_VERB_RE = re.compile(r"\[(ACK|CLAIM|RESULT|DECLINE|NO-ACTION)\s+([A-Za-z0-9_/-]+)\]")
_FALLBACK_TOOL_RE = re.compile(r"\[(TOOL|EXEC|DM)\s+(?:([a-zA-Z0-9_.-]+)\s+)?(\{([^{}]|\{[^{}]*\})*\})\]", re.S)
_FALLBACK_CONTRACT_FOOTER = (
"Reply: [ACK id] seen | [CLAIM id] mine | "
"[RESULT id] done | [DECLINE id] | [NO-ACTION id]."
)
def load_lookup_json(filename: str) -> Dict[str, Any]:
"""Load JSON from lookup_internal/ with mtime-based caching."""
target_path = LOOKUP_INTERNAL / filename
if not target_path.is_file():
return {}
try:
current_mtime = os.path.getmtime(target_path)
except OSError:
return _RAW_CACHE.get(filename, {})
if filename in _RAW_CACHE and _CACHE_MTIMES.get(filename) == current_mtime:
return _RAW_CACHE[filename]
try:
with open(target_path, "r", encoding="utf-8") as f:
data = json.load(f)
_RAW_CACHE[filename] = data
_CACHE_MTIMES[filename] = current_mtime
return data
except Exception as e:
print(f"[lookup_engine] Warning: Error reading {target_path}: {e}", file=sys.stderr)
return _RAW_CACHE.get(filename, {})
def _refresh_compiled_patterns_if_needed():
"""Checks regex_patterns.json mtime and recompiles if changed."""
global _COMPILED_PATTERNS
data = load_lookup_json("regex_patterns.json")
patterns_data = data.get("patterns", {})
target_path = LOOKUP_INTERNAL / "regex_patterns.json"
current_mtime = _CACHE_MTIMES.get("regex_patterns.json", 0.0)
compiled_mtime = _CACHE_MTIMES.get("_compiled_patterns_mtime", 0.0)
if current_mtime == compiled_mtime and _COMPILED_PATTERNS:
return
new_compiled = {}
for key, entry in patterns_data.items():
raw_pat = entry.get("pattern", "")
flag_names = entry.get("flags", [])
flags = 0
for fn in flag_names:
if hasattr(re, fn):
flags |= getattr(re, fn)
try:
new_compiled[key] = re.compile(raw_pat, flags)
except Exception as e:
print(f"[lookup_engine] Warning: Failed to compile pattern '{key}': {e}", file=sys.stderr)
_COMPILED_PATTERNS = new_compiled
_CACHE_MTIMES["_compiled_patterns_mtime"] = current_mtime
def get_compiled_pattern(name: str) -> Optional[re.Pattern]:
"""Retrieve a compiled pattern by name, hot-reloading if the database was modified."""
_refresh_compiled_patterns_if_needed()
return _COMPILED_PATTERNS.get(name)
def get_all_compiled_patterns() -> Dict[str, re.Pattern]:
"""Retrieve all compiled patterns with automatic hot-reloading."""
_refresh_compiled_patterns_if_needed()
return dict(_COMPILED_PATTERNS)
def get_result_regex() -> re.Pattern:
"""Return the canonical [RESULT ...] regex."""
p = get_compiled_pattern("result")
return p if p is not None else _FALLBACK_RESULT_RE
def get_verb_regex() -> re.Pattern:
"""Return the canonical [VERB ...] regex (ACK|CLAIM|RESULT|DECLINE|NO-ACTION)."""
p = get_compiled_pattern("verb")
return p if p is not None else _FALLBACK_VERB_RE
def get_tool_regex() -> re.Pattern:
"""Return the canonical [TOOL ...] regex."""
p = get_compiled_pattern("tool_call")
return p if p is not None else _FALLBACK_TOOL_RE
def get_contract_footer() -> str:
"""Return standard contract footer string."""
struct_data = load_lookup_json("sentence_structure.json")
cf = struct_data.get("structures", {}).get("contract_footer", {})
return cf.get("example") or _FALLBACK_CONTRACT_FOOTER
def parse_agent_utterance(text: str) -> List[Dict[str, Any]]:
"""Pass text through all registered patterns and extract matched tokens."""
_refresh_compiled_patterns_if_needed()
patterns_data = load_lookup_json("regex_patterns.json").get("patterns", {})
matches = []
for key, compiled in _COMPILED_PATTERNS.items():
m = compiled.search(text)
if m:
entry = patterns_data.get(key, {})
matches.append({
"pattern_key": key,
"pattern_name": entry.get("name", key),
"matched_text": m.group(0),
"named_groups": m.groupdict(),
"span": m.span()
})
return matches
def validate_outbound_sentence(text: str) -> Tuple[bool, Optional[str], Optional[str]]:
"""Pre-flight check for outbound messages sent via CLI.
Returns:
(is_valid, matched_kind, warning_or_hint)
"""
cleaned = text.strip()
# If message starts with bracketed protocol marker
if cleaned.startswith("["):
marker = cleaned.split("]")[0] + "]"
upper_marker = marker.upper()
if upper_marker.startswith("[WO:") or upper_marker.startswith("[WORKORDER:"):
wo_pat = get_compiled_pattern("work_order")
if wo_pat and not wo_pat.search(cleaned):
hint = (
"Notice: Message starts with a Work Order marker but does not match canonical structure.\n"
" Expected format: [WO:<id>] [from <sender>] <title> — <body>\n"
" Example: [WO:7fce46e0] [from super] Audit endpoints — Check GET /api/stats\n"
" Hint: Query 'box lookup sentence work_order' for full spec."
)
return False, "work_order", hint
return True, "work_order", None
if upper_marker.startswith("[ACK:") or upper_marker.startswith("[ACK "):
ack_pat = get_compiled_pattern("ack")
verb_pat = get_compiled_pattern("verb")
if (ack_pat and not ack_pat.search(cleaned)) and (verb_pat and not verb_pat.search(cleaned)):
hint = (
"Notice: Message looks like an ACK but deviates from standard syntax.\n"
" Expected format: [ACK:<id>] [from <sender>] or [ACK <id>]\n"
" Example: [ACK:7fce46e0] [from 646]"
)
return False, "ack", hint
return True, "ack", None
if upper_marker.startswith("[RESULT"):
res_pat = get_result_regex()
if not res_pat.search(cleaned):
hint = (
"Notice: Message starts with [RESULT] but deviates from standard syntax.\n"
" Expected format: [RESULT <job_id>] <status> <summary>\n"
" Example: [RESULT 7fce46e0] OK Task completed successfully"
)
return False, "result", hint
return True, "result", None
return True, "plain_message", None
-194
View File
@@ -1,194 +0,0 @@
#!/usr/bin/env python3
"""main-chat-watchdog.py — enforce the sidechat-first DM policy.
Tails dm-log.jsonl (read-only) and flags any DM actually SENT to main chat
without an explicit allow_main_chat marker, plus any placement_mismatch /
placement_failed events (message landed in the wrong chat). Distinguishes:
VIOLATION : type=sent, target=main, no tags.allow_main_chat -> policy broken (P0)
AUTHORIZED: type=sent, target=main, tags.allow_main_chat=true -> explicit opt-in
BLOCKED : type=main_chat_blocked -> policy working
MISMATCH : type=placement_mismatch -> placed in wrong chat (P0)
PFAILED : type=placement_failed -> placement hard-failed (P0)
State: byte-offset watermark in main-chat-watchdog.state (handles log
rotation via inode check). First run starts at EOF (no historical backfill —
old entries predate the allow_main_chat audit marker).
Violations are appended to logs/main-chat-violations.jsonl and printed to
stdout (journal). Exit 0 = clean, 1 = violations/P0s found, 2 = error.
Read-only: never modifies dm-log.jsonl.
"""
import json
import os
import sys
from datetime import datetime, timezone
BASE = "/home/super/Projects/NetVM"
DM_LOG = os.path.join(BASE, "dm-log.jsonl")
STATE_FILE = os.path.join(BASE, "main-chat-watchdog.state")
VIOLATIONS_LOG = os.path.join(BASE, "logs", "main-chat-violations.jsonl")
def utcnow():
return datetime.now(timezone.utc).isoformat()
def load_state():
try:
with open(STATE_FILE) as f:
return json.load(f)
except (FileNotFoundError, json.JSONDecodeError, ValueError):
return {}
def save_state(state):
tmp = STATE_FILE + ".tmp"
with open(tmp, "w") as f:
json.dump(state, f)
os.replace(tmp, STATE_FILE)
def main():
try:
st = os.stat(DM_LOG)
except FileNotFoundError:
print(f"watchdog ERROR: {DM_LOG} not found", file=sys.stderr)
return 2
state = load_state()
# First run (or rotation): start at EOF, don't backfill history that
# predates the allow_main_chat audit marker.
if state.get("inode") != st.st_ino:
offset = st.st_size
if state:
print(f"watchdog: log rotated or first run (inode {st.st_ino}), "
f"starting at EOF offset {offset}")
else:
offset = min(state.get("offset", 0), st.st_size)
violations = [] # P0: gate bypass (sent to main, no opt-in)
mismatches = [] # P0: placement_mismatch / placement_failed
blocked = 0
authorized = 0
scanned = 0
malformed = 0
try:
with open(DM_LOG, "r", encoding="utf-8", errors="replace") as f:
f.seek(offset)
for line in f:
line = line.strip()
if not line:
continue
scanned += 1
try:
ev = json.loads(line)
except json.JSONDecodeError:
malformed += 1
continue
etype = ev.get("type")
if etype == "main_chat_blocked":
# Policy working: the dm.py gate refused a main send.
blocked += 1
elif etype == "placement_mismatch":
# P0: verify-time URL check found the browser parked in a
# different chat than the intended target. The send
# completed but landed in the wrong place. Schema has two
# variants: sidechat ("expected_uuid") and main-drift
# ("expected": "main").
mismatches.append({
"ts": utcnow(),
"kind": "placement_mismatch",
"severity": "P0",
"dm_id": ev.get("id"),
"agent": ev.get("agent"),
"to": ev.get("to"),
"target": ev.get("target"),
"expected_uuid": ev.get("expected_uuid") or ev.get("expected"),
"actual_uuid": ev.get("actual_uuid"),
"attempt": ev.get("attempt"),
"event_ts": ev.get("ts"),
})
elif etype == "placement_failed":
# P0: the authoritative post-send placement check failed
# hard (sender exited 1, send NOT marked verified).
d = ev.get("detail") or {}
mismatches.append({
"ts": utcnow(),
"kind": "placement_failed",
"severity": "P0",
"dm_id": ev.get("id"),
"agent": ev.get("agent"),
"to": ev.get("to"),
"target": ev.get("target"),
"reason": ev.get("reason") or d.get("reason"),
"expected_uuid": ev.get("expected_uuid") or d.get("expected_uuid"),
"actual_url": ev.get("actual_url") or d.get("actual_url"),
"loop_attempt": ev.get("loop_attempt"),
"event_ts": ev.get("ts"),
})
elif etype == "sent" and ev.get("target") == "main":
tags = ev.get("tags") or {}
if tags.get("allow_main_chat"):
authorized += 1
else:
violations.append({
"ts": utcnow(),
"kind": "main_chat_send",
"severity": "P0",
"dm_id": ev.get("id"),
"agent": ev.get("agent"),
"to": ev.get("to"),
"target": "main",
"msg_preview": (ev.get("msg") or "")[:120],
"event_ts": ev.get("ts"),
})
new_offset = f.tell()
except OSError as e:
print(f"watchdog ERROR reading log: {e}", file=sys.stderr)
return 2
save_state({"offset": new_offset, "inode": st.st_ino})
findings = violations + mismatches
summary = (f"watchdog: scanned={scanned} blocked={blocked} "
f"authorized_main={authorized} "
f"violations={len(violations)} "
f"placement_mismatches={len(mismatches)} "
f"malformed={malformed}")
print(summary)
if findings:
try:
os.makedirs(os.path.dirname(VIOLATIONS_LOG), exist_ok=True)
with open(VIOLATIONS_LOG, "a", encoding="utf-8") as vf:
for v in findings:
vf.write(json.dumps(v) + "\n")
except OSError as e:
print(f"watchdog ERROR writing violations log: {e}", file=sys.stderr)
return 2
for v in violations:
print(f"VIOLATION main-chat send without opt-in: "
f"id={v['dm_id']} agent={v['agent']} to={v['to']} "
f"at={v['event_ts']} preview={v['msg_preview']!r}")
for m in mismatches:
if m["kind"] == "placement_mismatch":
print(f"P0 PLACEMENT_MISMATCH message landed in wrong chat: "
f"id={m['dm_id']} agent={m['agent']} to={m['to']} "
f"target={m['target']} expected={m['expected_uuid']} "
f"actual={m['actual_uuid']} at={m['event_ts']}")
else:
print(f"P0 PLACEMENT_FAILED placement check failed: "
f"id={m['dm_id']} agent={m['agent']} to={m['to']} "
f"target={m['target']} reason={m['reason']} "
f"expected={m['expected_uuid']} "
f"actual_url={m['actual_url']} at={m['event_ts']}")
return 1
return 0
if __name__ == "__main__":
sys.exit(main())
-205
View File
@@ -1,205 +0,0 @@
#!/usr/bin/env python3
"""Meta Accounts Center change-detection harness.
Captures a structural snapshot of the accountscenter.meta.com auth flow
via CDP inside a NetVM netns, diffs against the stored baseline.
Outcomes:
PASS - matches baseline (or first run establishes it)
CHANGED - structural diff detected; needs human review, baseline untouched
FAIL - automation itself broke (browser/CDP/network error)
Usage:
meta-ac-snapshot.py [--node NAME] [--promote] [--snapshot-dir DIR]
--node NetVM node to run in (default: phone)
--promote after human review, promote the latest snapshot to baseline
--snapshot-dir where snapshots live (default: ~/Projects/NetVM/snapshots/meta-ac)
"""
import argparse, base64, datetime, json, os, subprocess, sys, time
import urllib.parse, urllib.request
CDP_PORT = 19744
def log(*a):
print(*a, flush=True)
def ns_exec(node, cmd):
return subprocess.run(
["sudo", "-n", "ip", "netns", "exec", f"warp-{node}"] + cmd,
capture_output=True, text=True)
def norm_url(u):
"""Strip query/fragment — nonces change every visit."""
p = urllib.parse.urlparse(u)
return f"{p.scheme}://{p.host}{p.path}" if hasattr(p, 'host') else f"{p.scheme}://{p.hostname}{p.path}"
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--node", default="phone")
ap.add_argument("--promote", action="store_true")
ap.add_argument("--snapshot-dir", default=os.path.expanduser(
"~/Projects/NetVM/snapshots/meta-ac"))
args = ap.parse_args()
os.makedirs(args.snapshot_dir, exist_ok=True)
baseline_path = os.path.join(args.snapshot_dir, "baseline.json")
if args.promote:
snaps = sorted(f for f in os.listdir(args.snapshot_dir)
if f.startswith("snap-") and f.endswith(".json"))
if not snaps:
log("no snapshots to promote"); return 2
latest = os.path.join(args.snapshot_dir, snaps[-1])
data = json.load(open(latest))
data["promoted_at"] = datetime.datetime.now(datetime.timezone.utc).isoformat()
json.dump(data, open(baseline_path, "w"), indent=2)
log(f"promoted {snaps[-1]} -> baseline.json")
return 0
profile_dir = "/tmp/meta-ac-snap-profile"
subprocess.run(["rm", "-rf", profile_dir])
os.makedirs(profile_dir, exist_ok=True)
def http(path):
with urllib.request.urlopen(
f"http://127.0.0.1:{CDP_PORT}{path}", timeout=5) as r:
return json.loads(r.read())
# launch chromium inside the netns via a wrapper script
wrapper = "/tmp/meta-ac-snap-run.py"
open(wrapper, "w").write(WRAPPER_SRC)
log(f"launching chromium in warp-{args.node} (CDP {CDP_PORT})...")
proc = subprocess.Popen(
["sudo", "-n", "ip", "netns", "exec", f"warp-{args.node}",
"python3", wrapper, str(CDP_PORT), profile_dir],
stdout=subprocess.PIPE, stderr=subprocess.STDOUT, text=True)
try:
out, _ = proc.communicate(timeout=120)
except subprocess.TimeoutExpired:
proc.kill(); log("FAIL: harness timed out"); return 1
print(out)
# wrapper prints SNAPSHOT_JSON=<json> on success
snap = None
for line in out.splitlines():
if line.startswith("SNAPSHOT_JSON="):
snap = json.loads(line[len("SNAPSHOT_JSON="):])
if not snap:
log("FAIL: no snapshot captured"); return 1
snap["node"] = args.node
snap["captured_at"] = datetime.datetime.now(datetime.timezone.utc).isoformat()
# egress ip for context
try:
r = ns_exec(args.node, ["curl", "-s", "--max-time", "8",
"https://api.ipify.org"])
snap["egress_ip"] = r.stdout.strip()
except Exception:
snap["egress_ip"] = "unknown"
ts = datetime.datetime.now(datetime.timezone.utc).strftime("%Y%m%d-%H%M%S")
snap_path = os.path.join(args.snapshot_dir, f"snap-{ts}.json")
json.dump(snap, open(snap_path, "w"), indent=2)
log(f"snapshot saved: {snap_path}")
if not os.path.exists(baseline_path):
json.dump(snap, open(baseline_path, "w"), indent=2)
log("PASS: baseline established (first run)")
return 0
baseline = json.load(open(baseline_path))
diffs = diff_snapshots(baseline, snap)
if not diffs:
log("PASS: matches baseline")
return 0
log("CHANGED: structural diff detected (baseline untouched):")
for d in diffs:
log(f" - {d}")
log("review with: diff baseline.json snap-<ts>.json")
log("promote after review with: --promote")
return 3
def diff_snapshots(base, snap):
diffs = []
b_chain = [norm_url(u) for u in base.get("redirect_chain", [])]
s_chain = [norm_url(u) for u in snap.get("redirect_chain", [])]
if b_chain != s_chain:
diffs.append(f"redirect_chain changed: {b_chain} -> {s_chain}")
for key in ("forms", "inputs", "buttons"):
b = sorted(base.get("dom_markers", {}).get(key, []))
s = sorted(snap.get("dom_markers", {}).get(key, []))
if b != s:
added = [x for x in s if x not in b]
removed = [x for x in b if x not in s]
diffs.append(f"dom_markers.{key}: added={added} removed={removed}")
if base.get("final_title") != snap.get("final_title"):
diffs.append(f"final_title: {base.get('final_title')!r} -> {snap.get('final_title')!r}")
return diffs
WRAPPER_SRC = '''
import json, subprocess, sys, time, os, urllib.request, base64
CDP_PORT = int(sys.argv[1])
PROFILE_DIR = sys.argv[2]
def http(path):
with urllib.request.urlopen(f"http://127.0.0.1:{CDP_PORT}{path}", timeout=5) as r:
return json.loads(r.read())
logf = open("/tmp/meta-ac-snap-chrome.log", "w")
proc = subprocess.Popen(["chromium", "--headless=new", "--disable-gpu", "--no-sandbox",
"--disable-dev-shm-usage", f"--user-data-dir={PROFILE_DIR}",
f"--remote-debugging-port={CDP_PORT}", "--remote-allow-origins=*", "about:blank"],
stdout=logf, stderr=subprocess.STDOUT)
try:
for i in range(30):
try:
ver = http("/json/version")
if "webSocketDebuggerUrl" in ver: break
except Exception: pass
time.sleep(1)
else:
print("FAIL: CDP never came up"); sys.exit(1)
import websocket
bws = websocket.create_connection(ver["webSocketDebuggerUrl"], timeout=20)
bws.send(json.dumps({"id": 1, "method": "Target.createTarget",
"params": {"url": "https://accountscenter.meta.com"}}))
target_id = json.loads(bws.recv())["result"]["targetId"]
bws.close()
# redirect chain: seed with the navigation target (we always start
# there), then poll for where Meta sends us. Seeding fixes the race
# where a fast redirect is missed by the poll interval.
START_URL = "https://accountscenter.meta.com/"
chain, seen = [START_URL], {START_URL}
for _ in range(24):
time.sleep(2)
for t in http("/json/list"):
if t.get("id") == target_id or "meta.com" in t.get("url", ""):
u = t["url"]
if u not in seen:
seen.add(u); chain.append(u)
title = t.get("title", "")
break
# dom markers from the final tab
tab_ws = None
for t in http("/json/list"):
if t.get("id") == target_id or "meta.com" in t.get("url", ""):
tab_ws = t["webSocketDebuggerUrl"]; final_url = t["url"]; break
ws = websocket.create_connection(tab_ws, timeout=20)
js = """JSON.stringify({
forms: [...document.forms].map(f => f.id || f.name || '(anon)'),
inputs: [...document.querySelectorAll('input')].map(i => i.name || i.type || '(anon)'),
buttons: [...document.querySelectorAll('button, [role=button]')].map(b => (b.innerText||'').trim()).filter(Boolean)
})"""
ws.send(json.dumps({"id": 1, "method": "Runtime.evaluate",
"params": {"expression": js, "returnByValue": True}}))
markers = json.loads(json.loads(ws.recv())["result"]["result"]["value"])
# dedupe buttons, keep order
markers["buttons"] = list(dict.fromkeys(markers["buttons"]))
ws.close()
snap = {"redirect_chain": chain, "final_url": final_url,
"final_title": title, "dom_markers": markers}
print("SNAPSHOT_JSON=" + json.dumps(snap))
finally:
proc.terminate()
'''
if __name__ == "__main__":
sys.exit(main())
-218
View File
@@ -1,218 +0,0 @@
#!/usr/bin/env python3
"""Meta Accounts Center read API (via authenticated browser session).
Usage:
meta-acct.py list-linked <agent> # linked profiles in this Meta Account
meta-acct.py security-status <agent> # login & recovery / security checkup summary
meta-acct.py login-activity <agent> # "Where you're logged in" sessions
The agent's browser must have an active Meta session (phone/email OTP login).
Reads NODES.md for the CDP port. Outputs JSON.
Scope: Accounts Center linkage/security surface only. No writes.
"""
import json, sys, time, urllib.request, os
NETVM = os.environ.get("NETVM_DIR")
if not NETVM:
script_dir = os.path.dirname(os.path.abspath(__file__))
parent = os.path.dirname(script_dir)
if os.path.exists(os.path.join(parent, "NODES.md")):
NETVM = parent
else:
NETVM = "/home/super/Projects/NetVM"
AC_BASE = "https://accountscenter.meta.com"
def cdp_port(agent):
for line in open(os.path.join(NETVM, "NODES.md")):
line = line.strip()
if not line.startswith("|"):
continue
cells = [c.strip() for c in line.strip("|").split("|")]
if len(cells) >= 4 and cells[0] == agent:
return int(cells[3])
raise SystemExit(f"meta-acct: unknown agent '{agent}' (not in NODES.md)")
def cdp_get(port, path, method="GET", timeout=10):
req = urllib.request.Request(f"http://127.0.0.1:{port}{path}", method=method)
with urllib.request.urlopen(req, timeout=timeout) as r:
return json.loads(r.read())
def get_ac_tab(port):
tabs = cdp_get(port, "/json/list")
for t in tabs:
if "accountscenter.meta.com" in t.get("url", "") and t.get("type") == "page":
return t
# create one
return cdp_get(port, f"/json/new?{AC_BASE}/", method="PUT")
def evaluate(port, ws_url, js, timeout=15):
import websocket
ws = websocket.create_connection(ws_url, timeout=timeout)
try:
ws.send(json.dumps({"id": 1, "method": "Page.navigate",
"params": {"url": js[0]}}))
ws.recv()
time.sleep(4)
ws.send(json.dumps({"id": 2, "method": "Runtime.evaluate",
"params": {"expression": js[1], "returnByValue": True}}))
resp = json.loads(ws.recv())
return json.loads(resp["result"]["result"]["value"])
finally:
ws.close()
def cmd_list_linked(port):
tab = get_ac_tab(port)
data = evaluate(port, tab["webSocketDebuggerUrl"], (
f"{AC_BASE}/account_overview/",
"""JSON.stringify((() => {
const h2 = [...document.querySelectorAll('h2')]
.find(e=>e.innerText.trim()==='Profiles');
// Full section is 3 levels up (H2 > DIV > DIV > MAIN)
const container = h2?.parentElement?.parentElement?.parentElement;
return {
email: (document.body.innerText.match(
/[\\w.+-]+@[\\w-]+\\.[\\w.]+/)||[])[0]||null,
section: container?.innerText?.slice(0,1000) || null
};
})())"""))
# Parse: lines after "Profiles" are name/type pairs until "Add more",
# then "AI and devices" section follows
profiles = []
section = data.get("section", "")
if section:
lines = [l.strip() for l in section.split("\n") if l.strip()]
try:
start = lines.index("Profiles") + 1
except ValueError:
start = 0
i = start
while i < len(lines):
if lines[i] == "Add more":
i += 1
continue
if lines[i] == "AI and devices":
# Next line is the device/service name
if i + 1 < len(lines):
profiles.append({"name": lines[i+1], "type": "ai_device"})
break
# Expect name + type pair
if i + 1 < len(lines) and lines[i+1] not in (
"Add more", "AI and devices", "Profiles"):
profiles.append({"name": lines[i],
"type": lines[i+1].lower()})
i += 2
else:
i += 1
return {"email": data.get("email"), "profiles": profiles}
def cmd_security_status(port):
tab = get_ac_tab(port)
data = evaluate(port, tab["webSocketDebuggerUrl"], (
f"{AC_BASE}/password_and_security/",
"""JSON.stringify({
checkup: document.body.innerText.match(
/Meta Security Checkup\\s*(\\d+) recommended actions/)?.[1] || null,
sections: [...document.querySelectorAll('h2')]
.map(e=>e.innerText.trim()).slice(0,10),
body: document.body.innerText.slice(0,2000)
})"""))
return {"security_checkup_actions": data.get("checkup"),
"sections": data.get("sections")}
def cmd_login_activity(port):
tab = get_ac_tab(port)
# "Where you're logged in" is on the password_and_security page
data = evaluate(port, tab["webSocketDebuggerUrl"], (
f"{AC_BASE}/password_and_security/",
"""JSON.stringify({
body: document.body.innerText.slice(0,4000)
})"""))
body = data.get("body", "")
# Extract the "Where you're logged in" section
idx = body.find("Where you're logged in")
section = body[idx:idx+1500] if idx >= 0 else ""
return {"where_logged_in": section}
def cmd_link_instagram(agent, port):
"""Automate Meta Accounts Center linking flow and second-click OAuth handoff."""
import websocket
tab = get_ac_tab(port)
ws = websocket.create_connection(tab["webSocketDebuggerUrl"], timeout=15)
try:
# Step 1: Navigate to manage accounts
ws.send(json.dumps({"id": 1, "method": "Page.navigate", "params": {"url": f"{AC_BASE}/manage/"}}))
ws.recv()
time.sleep(3)
# Step 2: Look for 'Add profiles and devices' button or check if already linked
expr_find_add = """(() => {
const btns = [...document.querySelectorAll("div[role='button'], button, a")];
const addBtn = btns.find(b => /add profiles|add accounts/i.test((b.innerText||"").trim()));
if (addBtn) {
addBtn.click();
return {status: "clicked_add", text: addBtn.innerText};
}
return {status: "no_add_button"};
})()"""
ws.send(json.dumps({"id": 2, "method": "Runtime.evaluate", "params": {"expression": expr_find_add, "returnByValue": True}}))
res2 = json.loads(ws.recv()).get("result", {}).get("result", {}).get("value", {})
time.sleep(3)
# Step 3: Check dialog / popup for Instagram option or FXCAL handoff
expr_handle_dialog = """(() => {
const btns = [...document.querySelectorAll("div[role='button'], button, a")];
// Check for Add Instagram button or Continue/Confirm
const igBtn = btns.find(b => /instagram/i.test((b.innerText||"").trim()) && /add|connect/i.test((b.innerText||"").trim()));
if (igBtn) {
igBtn.click();
return {status: "clicked_instagram_option"};
}
const confirmBtn = btns.find(b => /continue|confirm|yes, finish/i.test((b.innerText||"").trim()));
if (confirmBtn) {
confirmBtn.click();
return {status: "clicked_confirm", text: confirmBtn.innerText};
}
return {status: "dialog_scanned", current_url: window.location.href};
})()"""
ws.send(json.dumps({"id": 3, "method": "Runtime.evaluate", "params": {"expression": expr_handle_dialog, "returnByValue": True}}))
res3 = json.loads(ws.recv()).get("result", {}).get("result", {}).get("value", {})
time.sleep(2)
return {
"status": "success",
"step1_add": res2,
"step2_dialog": res3,
"current_url": tab.get("url")
}
finally:
ws.close()
def main():
if len(sys.argv) < 3:
print(__doc__.strip().split("\n")[0])
print("Usage: meta-acct.py <list-linked|security-status|login-activity|link-instagram> <agent>")
sys.exit(2)
cmd, agent = sys.argv[1], sys.argv[2]
port = cdp_port(agent)
try:
if cmd == "list-linked":
out = cmd_list_linked(port)
elif cmd == "security-status":
out = cmd_security_status(port)
elif cmd == "login-activity":
out = cmd_login_activity(port)
elif cmd == "link-instagram":
out = cmd_link_instagram(agent, port)
else:
raise SystemExit(f"meta-acct: unknown command '{cmd}'")
except Exception as e:
print(json.dumps({"error": str(e), "agent": agent}))
sys.exit(1)
out["agent"] = agent
out["checked_at"] = int(time.time())
print(json.dumps(out, indent=2))
if __name__ == "__main__":
main()

Some files were not shown because too many files have changed in this diff Show More