Files
box/bin/swarm_worker/mainloop-bridge.md

4.8 KiB

Swarm Worker ↔ Main Loop Bridge — Design Doc

Bridge Builder 5/5 | 2026-10-05 | Read-only design (no changes to self_main_loop.py)

Context

Sibling builders produced a tmux-hosted worker daemon for bl:

  • executor.py — sandboxed task execution (shell vs prose heuristic, refuse-list)
  • poller.py — slot claiming
  • reporter.py — result recording via box swarm-report <id> <slot> (direct, reliable)

The user's directive: "we should be answering our own main chat via our main loop." This doc designs how swarm work becomes visible through self_main_loop.py without modifying it.

How the main loop works (verified against live code)

bin/self_main_loop.py (995 lines) is a timer-driven read-side loop:

  1. Reads each watched agent's muse.ai main chat + registered sidechats (via muse_hybrid gateway, box-chat.py fallback).
  2. On new activity, composes a digest and posts it to the agent's prompting sidechat via dm.py send (send_prompt()).
  3. Actionable digests (containing (?) or (!)) get [JOB ml-<agent>-<ts>]
    • contract footer → --expect-reply → followup record. Informational digests go via fast gateway, no followup.
  4. Response verbs [ACK|CLAIM|RESULT|DECLINE|NO-ACTION <id>] are matched by response-harvester.py to resolve followups.

Key detail — check_sidechats() skips individual messages containing "[JOB ", "Heartbeat check", or "[from:super]". It does NOT skip entire sidechats. A [RESULT ...] or [SWARM-DONE ...] post is not filtered.

How swarm slot DMs arrive (verified against dm-log.jsonl)

Schema: {type, id, agent, to, target, msg, tags, ts}. Lifecycle: send_start → sidechat_autoprovisioned → verified → sent.

  • Dispatch: dm.py send --agent <dispatcher> --to <worker> --target sw-<id>-s<slot> with body [JOB sw-<id>/<slot>] Task for swarm slot N:\n<task>\n\nReply with [RESULT sw-...]
  • Target auto-provisions a sidechat; 23 slot sidechats are registered in job-sidechats.json with per-slot agent attribution (e.g. agent: "dev").
  • Harvester (response-harvester.py:773) bridges sidechat [RESULT <swarm_id>/<slot>] markers into box swarm-report — but the sibling's reporter.py calls swarm-report directly (no DM round-trip).

Recommendation: (a) direct sidechat post, main loop as visibility layer

Result recording (unchanged): worker → box swarm-report directly (sibling's reporter.py). Reliable, no harvester-scan dependency.

Main-loop bridge (new, zero code changes): after recording the result, the worker posts a short human-readable completion note to the slot's own sidechat via dm.py send — deliberately without a [JOB tag so the main loop's check_sidechats() does not filter it:

[SWARM-DONE sw-20261005-151307-a222/0] OK in 12.4s: <first ~200 chars of output>

Why this works with no modifications:

  1. Slot sidechats are already registered in job-sidechats.json with agent attribution → get_monitored_sidechats() includes them for that agent.
  2. The note contains no [JOB → not skipped by the message filter.
  3. Next main-loop tick picks it up → digest → agent's prompting sidechat → the agent sees the completion in the same surface as everything else. This is "answering via the main loop": completions surface through the existing digest machinery instead of a parallel notification channel.

Why not (b) queue via send_prompt(): send_prompt() is built for main-chat digests → prompting sidechats, with [JOB ml-...] followup semantics. Swarm results are not digests; routing them through would add timer latency, misuse the ml- followup namespace, and conflate two concerns.

Why not (c) worker as main-loop participant: the main loop monitors muse.ai agent accounts (muse_hybrid.get_history(agent, ...)). The worker is a bl daemon with no agent account. Participation-by-proxy (writing to monitored sidechats) achieves the same visibility with no identity hacks.

Watchdog (no change needed, documented here)

check_sidechats() also gives stall detection for free: a claimed slot whose sidechat shows no [SWARM-DONE]/[RESULT] within the slot timeout is visible as inactivity in the digest. A future enhancement (not this bridge) could add an explicit stall-threshold alert.

Interface for the worker daemon

After reporter.post_result(...) succeeds, call: notify_via_mainloop(swarm_id, slot_index, summary_text) (prototype: bin/swarm_worker/mainloop_notify.py). It resolves sw-<swarm_id>-s<slot_index> → sidechat UUID via dm.resolve_sidechat_target, then dm.py send with the [SWARM-DONE ...] note. Dry-run mode posts nothing.

Files

  • This doc: bin/swarm_worker/mainloop-bridge.md
  • Prototype: bin/swarm_worker/mainloop_notify.py (dry-run safe, never auto-wired)