Files
box/docs/LOOP-MANAGEMENT.md
T

12 KiB

NetVM Intrinsic Loop Management — Operational Runbook & Architecture Specification

Box is the main surface. All operator work goes through Box (box.muse-dev.online). The web UI, box CLI, and agents share the same API endpoints. No UI-only powers.

1. Executive Overview

NetVM coordinates an autonomous agent mesh (muse, pip, 646, opm, super) communicating via inter-agent direct messages (DMs), scheduled jobs, and live browser sidechats. Intrinsic Loops represent communication cycles requiring closure (e.g. follow-ups, results, acknowledgements).

To prevent silent failures, stale deadlines, or rogue infinite nudging, NetVM provides External Loop Management:

  • Dual-Surface Architecture: Real-time local CLI management on bl (super and box commands) synchronized with an operator Web Console on the Google Cloud VM (https://box.muse-dev.online/).
  • Dynamic Runtime Control Variables: Typed runtime knobs controlling sampling cadences, silence thresholds, and retry policies with atomic rollbacks.
  • Hierarchical Modulation: Rule cascade determining follow-up tracking policies scoped by (input_type, subtype, agent).
  • Progressive Auto-Remediation: Background daemon healing soft breaks while loudly escalating hard breaks.

2. Architecture Diagram

flowchart TD
    subgraph VM ["Google Cloud Gateway VM (34.139.37.135)"]
        UI["Box Web Console (/srv/box/www)"]
        Board["board.service (/srv/board/server.py)"]
        UI -->|HTTP /api/box/loop/*| Board
    end

    subgraph Tailnet ["Tailscale Secure Mesh (100.123.153.75)"]
        Board -->|SSH Allowlisted RPC| BoxCtl["bin/box-ctl.py"]
    end

    subgraph BL ["Local Management Node (bl)"]
        BoxCtl --> VarEng["Variables Engine (bin/variables.py)"]
        BoxCtl --> ModEng["Modulation Strategy (bin/modulate.py)"]
        BoxCtl --> GravEng["Loop Diagnostics (bin/gravity.py)"]

        CLI["CLI Orchestrator (bin/super-cli.py)"]
        CLI --> VarEng
        CLI --> ModEng
        CLI --> GravEng

        Daemon["systemd: loop-remediator.timer (15m)"]
        Daemon --> GravEng

        GravEng -->|Soft Heal| Followups["followups.json"]
        GravEng -->|Hard Break Alert| DMLog["bin/dm.py -> opm"]
        GravEng -->|Audit Trail| JobLog["job-log.jsonl"]
    end

3. Dual-Surface API & CLI Reference

3.1. Runtime Control Variables

The runtime variables engine (bin/variables.py) enforces type constraints, ranges, and dual-sync persistence between /srv/box/variables.json and ./variables.json. Every mutation is appended to variables-history.jsonl.

CLI Commands

# List all registered variables, values, units, and ranges
super vars list
# or
box vars-list

# Get specific variable
super vars get loop_health_threshold

# Set a variable (validated against schema)
super vars set loop_health_threshold 0.65

# Reset variable to default
super vars reset loop_health_threshold

# Inspect audit history
super vars history [name] [limit]

# Atomic rollback
super vars rollback loop_health_threshold

VM REST Endpoints (/api/box/loop/*)

  • GET /api/box/loop/vars → Retrieves full dictionary of runtime variables.
  • POST /api/box/loop/vars → Body: {"name": "...", "value": ...} (returns 202 Accepted).
  • POST /api/box/loop/vars/reset → Body: {"name": "..."}.
  • GET /api/box/loop/vars/history?name=... → Returns append-only revision history.

3.2. Hierarchical Modulation Strategy

Follow-up tracking behavior (bin/modulate.py) is resolved hierarchically across four precedence levels down to the builtin table:

\text{Override Precedence: } (T, S, A) \succ (T, \text{None}, A) \succ (T, S, \text{None}) \succ (T, \text{None}, \text{None}) \succ \text{Builtin}
  1. Exact match: (input_type, subtype, agent)
  2. Agent default: (input_type, None, agent)
  3. Subtype default: (input_type, subtype, None)
  4. Type default: (input_type, None, None)
  5. Builtin table fallback

CLI Commands

# Show modulation matrix (builtins + active overrides)
super strat show

# Set override
super strat set manual --timeout 1800 --nudges 1

# Set agent-specific override
super strat set manual --agent pip --no-track

# Reset override
super strat reset manual --agent pip

VM REST Endpoints

  • GET /api/box/loop/strat → Returns merged modulation matrix.
  • POST /api/box/loop/strat → Body: {"input_type": "...", "subtype": "...", "agent": "...", ...}.
  • POST /api/box/loop/strat/reset → Resets override for key.

3.3. Loop Diagnostics & Progressive Remediation

Loop health (bin/gravity.py) tracks loop status across agents:

\text{Health Ratio} = \frac{\text{Closed} + \text{Answered}}{\text{Landed}}

Progressive Remediation Workflow

  1. Soft Breaks (Auto-Healed):
    • Answered Loops: If a pending follow-up in followups.json has a matching reply detected in dm-log.jsonl, it is automatically marked resolved with note auto-healed: reply detected in dm-log.
    • Expired Nudges: If a loop deadline has lapsed but allowable nudges remain, the deadline is updated to now and bin/followup-sweeper.py is invoked immediately.
  2. Hard Breaks (Loudly Escalated):
    • silent_agent: Agent unresponsive after exhausting all allowed nudges.
    • auth_rot: Missing or corrupted SSH Ed25519 signing key (~/.ssh/id_ed25519).
    • scheduler_death: Systemd user session or timer infrastructure offline.
    • Escalation Actions: Emits structured event to job-log.jsonl and dispatches an immediate DM alert to opm on main.

CLI Commands

# View fleet loop health table and ratio
super loop health

# View all active / reconstructed loops
super loop status --limit 50

# Diagnose detected loop breakages
super loop breaks

# Manually resolve a stuck loop
super loop close <dm_id> "Resolved via operator intervention"

# Trigger manual remediation pass
super loop remediate [--dry-run]

VM REST Endpoints

  • GET /api/box/loop/health → JSON summary of fleet ratios and health verdicts.
  • GET /api/box/loop/status?limit=50 → Active loop instances.
  • POST /api/box/loop/resolve → Body: {"dm_id": "...", "note": "..."}.
  • POST /api/box/loop/remediate → Runs progressive remediation cycle.

Followup record fields & nudge→reply matching (2026-10-04 fix set)

Followup records in followups.json carry two fields introduced by the 2026-10-04 reliability fix set (see docs/SIDECHAT-RELIABILITY.md):

  • thread_uuid backfill — when a followup is created but the sidechat thread autoprovisioning failed (no UUID captured), the sweeper retries on each nudge. Once a nudge send succeeds with autoprovisioning (sidechat_autoprovisioned in dm-log.jsonl), the sweeper backfills thread_uuid into the record. Once set, the field is never overwritten with null — a followup with a null thread_uuid can never match a reply, so it would ghost-nudge then escalate on a thread that never existed.
  • final_nudge_target — set to "main" when the final nudge routes to the main chat instead of a sidechat (routing policy: nudge in-thread first, main only on the final attempt).

Nudge→reply matching matrix (implemented in bin/response-harvester.py):

Nudge routing Reply location Resolves when
In-thread nudge Assistant reply in the thread whose UUID == thread_uuid UUID match
Final nudge → main Assistant reply in main chat final_nudge_target == "main"
Any Reply in a different thread Never auto-resolves (manual loop close)

The sweeper also backfills thread_uuid immediately after a successful nudge send when the send itself autoprovisioned the thread, so later replies match.

dm.py placement assertion — before sending to a sidechat, dm.py fails closed (loud refusal, no verified:true) if the post-navigation browser URL does not contain the target thread UUID. The "main" target is exempt. This catches silent sidechat use failures and wrong-thread sends.

Multi-[RESULT] harvesting — the harvester processes every [RESULT …] marker in a message (previously only the first), so replies that batch several job results each resolve their followup and trigger their own chain step.


4. Background Services & Daemons

On node bl, loop remediation is managed by systemd user units:

Inspect service status:

systemctl --user status loop-remediator.timer
journalctl --user -u loop-remediator.service -n 20 --no-pager

5. Operator Troubleshooting Runbook

Incident A: Fleet Health Drops Below Threshold (< 50%)

  1. Run super loop health to pinpoint the offending agent node.
  2. Run super loop breaks to see whether loops are NUDGED, ESCALATED, or BROKEN.
  3. If an agent is unresponsive:
    • Check container process: super fleet status.
    • Send diagnostic ping: super dm send --to <agent> --target "<agent tasks>" "Liveness check".
  4. Run super loop remediate to auto-heal any lagged answer states.

Incident B: Web Console Mutations Fail (403 or 500)

  1. Verify operator authentication: Ensure valid PIN session cookie or Bearer token on https://box.muse-dev.online/.
  2. Verify Tailnet SSH bridge:
    • From VM: ssh super@100.123.153.75 /home/super/Projects/NetVM/bin/box-ctl.py loop-health.
    • Check box-ctl.jsonl on bl for allowlisted action audit records.
  3. Check board.service logs on VM: sudo journalctl -u board -n 50 --no-pager.

Incident C: Accidental Variable Corruption

  1. View audit history: super vars history <variable_name>.
  2. Rollback to prior known good value: super vars rollback <variable_name>.
  3. If necessary, reset to hardcoded schema default: super vars reset <variable_name>.

6. Verification & Automated Testing

All operational modules are covered by the comprehensive unit test suite in tests/:

# Run complete test suite (26 passing tests)
python3 -m unittest discover -s tests -v