Files
box/docs/AGENT-TOOLING.md
T

6.5 KiB

Agent Tooling & Subagent Delegation Guide

Box is the main surface. All operator work goes through Box (box.muse-dev.online). The web UI, box CLI, and agents share the same API endpoints. No UI-only powers.

Welcome, operator. The NetVM environment provides you with the unified box command line tool (/usr/local/bin/box) for executing tasks, spawning sub-agents, and communicating with peers across the fleet.


1. Spawning Sub-Agents (box deploy subagent)

When you receive a complex task, large audit, or background verification, prioritize delegating sub-components to an autonomous sub-agent.

box deploy subagent --agent <self> --title "<task-name>" "<prompt>"

Parameters:

  • --agent: Your own identity (646, pip, opm, or muse).
  • --title: Brief descriptive title for the sub-agent session.
  • prompt: The specific instructions and criteria for the sub-agent.

Example:

box deploy subagent --agent 646 --title "exec-api-audit" "Audit the op-based allowlist on exec.muse-dev.online and report back valid ops."

2. Cross-Operator DMs & Relay (box dm send)

To coordinate with your peer operators (opm, 646, pip, muse) or hand off tasks across sidechats:

box dm send --agent <self> --to <recipient> --target <sidechat> "<message>"

Target Sidechats:

  • 646-pip: Bilateral coordination thread between 646 and Pip.
  • 646-opm-coord: Coordination thread between 646 and OPM.
  • <agent> tasks: Dedicated task sidechat for an individual agent (e.g. 646 tasks, pip tasks).

Example:

box dm send --agent 646 --to pip --target 646-pip "Hey Pip, start-page onboarding review draft is ready. Please confirm when ready to review."

3. Direct Gateway Tooling (muse & box muse)

You can directly interact with the headless Muse gateway inside your isolated network namespace using either muse or box muse:

# Using native muse wrapper (interactive prompt & account enforcement)
muse -a <account> chat               # Interactive conversational REPL with thread selection
muse -a <account> chat --thread <id> # Direct conversational REPL in specified thread
muse -a <account> status
muse -a <account> threads
muse -a <account> history --thread <thread_uuid> --limit 10
muse -a <account> send --thread <thread_uuid> "<message>"

# If invoked without -a/--account, it displays valid accounts and usage instructions:
muse

# Alternatively via box CLI:
box muse <self> threads
box muse <self> history --thread <thread_uuid> --limit 10
box muse <self> unread
box muse <self> session-start --title "<title>"
box muse <self> send --thread <thread_uuid> "<message>"

4. Shared Tmux Socket Tooling (muse tmux & [TOOL tmux.*])

Agents and operators can spawn background sessions and send keystrokes to long-running tasks via the shared socket /tmp/tmux-muse.sock. All session stdout/scrollback is automatically piped and persisted to logs/tmux/<session>.log for auditing and post-mortem analysis.

Socket Architecture & Isolation

  • Agent Socket (/tmp/tmux-muse.sock): Exclusively reserved for agent execution, automated work orders, and operator inspections of agent tasks.
  • Operator Socket (/tmp/tmux-1000/default): Reserved for user desktop sessions (main, etc.).
  • Server Persistence Hardening: The operator tmux server runs with set -s exit-empty off and set -s exit-unattached off so background sessions persist when clients disconnect or windows close.
  • Inactivity TTL: Inactive unattached sessions are automatically pruned after 2 hours (120 minutes) by bin/netvm-reaper.sh or via explicit pruning.

Via Native CLI:

# List sessions on shared socket
muse <account> tmux list

# Create or send commands to a background session
muse <account> tmux send <session> "command to execute"

# Capture recent scrollback output from a session
muse <account> tmux capture <session> --lines 30

# Kill a session
muse <account> tmux kill <session>

# Prune unattached sessions inactive for >2 hours (or custom --ttl in seconds)
muse <account> tmux prune [--ttl 7200]

Via Autonomous Agent Directives:

Agents can emit structured tool calls in sidechats:

[TOOL tmux.new {"session": "build-worker", "command": "python3 /srv/worker.py"}]
[TOOL tmux.send {"session": "build-worker", "keys": "ls -la"}]
[TOOL tmux.capture {"session": "build-worker", "lines": 20}]
[TOOL tmux.list {}]
[TOOL tmux.prune {"ttl": 7200}]
[TOOL tmux.kill {"session": "build-worker"}]

5. Work Orders ([WO:...]) & Prompt Envelope Specification

When jobs are dispatched to agents via bin/job-dispatch.py, they are wrapped in an actionable Work Order envelope generated by bin/prompt_envelope.py.

Work Order Structure

  • Header: Prefixed with [WO:<wo_id>] WORK ORDER - ACTION REQUIRED, NOT INFORMATIONAL.
    • <wo_id> is a deterministic 8-character identifier derived from the job ID.
  • Background Session: Automatically sets up a dedicated tmux session on the shared socket: work-<agent>-<wo_id>.
  • Immediate Tool Directives: The envelope enforces immediate execution rather than dry prose by specifying the opening tool calls:
    1. [TOOL tmux.new {"session": "work-<agent>-<wo_id>", "command": "bash"}]
    2. [TOOL tmux.send {"session": "work-<agent>-<wo_id>", "keys": "..."}]
    3. Native subagent spawn or cron timer directive.
  • Clean Runtime Context: Includes THREAD, JOB, and AGENT identity parameters. Container-inaccessible host SSH key paths are stripped to ensure agents never enter auth refusal loops.
  • Completion Contract: When the task execution finishes, the agent concludes the reply with:
    [RESULT <job_id>] <one-line summary of what ran and completed>
    
    The harvester (bin/response-harvester.py) detects this line and records the job outcome.

6. Best Practices & Invariants

  1. Sidechat-First Policy: All inter-agent coordination, subagent tasks, and heartbeats must stay in sidechats. Do not send automated routine messages to main chat (see CHAT_POLICY.md).
  2. Sub-Agent Prioritization: Break down complex diagnostic or verification jobs by delegating sub-tasks to dedicated subagent threads.
  3. Execution Reality: Work is only real if tool calls ran. Never provide purely verbal confirmation for tasks requiring system inspection or execution.
  4. Attribution & Result Tagging: For scheduled jobs and work orders, always conclude your response with [RESULT <job_id>] <summary>.