Files
box/bin/swarm_worker/mainloop-bridge.md
T

96 lines
4.8 KiB
Markdown
Raw Normal View History

# Swarm Worker ↔ Main Loop Bridge — Design Doc
**Bridge Builder 5/5 | 2026-10-05 | Read-only design (no changes to self_main_loop.py)**
## Context
Sibling builders produced a tmux-hosted worker daemon for bl:
- `executor.py` — sandboxed task execution (shell vs prose heuristic, refuse-list)
- `poller.py` — slot claiming
- `reporter.py` — result recording via `box swarm-report <id> <slot>` (direct, reliable)
The user's directive: *"we should be answering our own main chat via our main loop."*
This doc designs how swarm work becomes visible through `self_main_loop.py`
without modifying it.
## How the main loop works (verified against live code)
`bin/self_main_loop.py` (995 lines) is a timer-driven read-side loop:
1. Reads each watched agent's muse.ai **main chat** + registered **sidechats**
(via `muse_hybrid` gateway, `box-chat.py` fallback).
2. On new activity, composes a digest and posts it to the agent's **prompting
sidechat** via `dm.py send` (`send_prompt()`).
3. Actionable digests (containing `(?)` or `(!)`) get `[JOB ml-<agent>-<ts>]`
+ contract footer → `--expect-reply` → followup record. Informational
digests go via fast gateway, no followup.
4. Response verbs `[ACK|CLAIM|RESULT|DECLINE|NO-ACTION <id>]` are matched by
`response-harvester.py` to resolve followups.
Key detail — `check_sidechats()` **skips individual messages** containing
`"[JOB "`, `"Heartbeat check"`, or `"[from:super]"`. It does NOT skip entire
sidechats. A `[RESULT ...]` or `[SWARM-DONE ...]` post is *not* filtered.
## How swarm slot DMs arrive (verified against dm-log.jsonl)
Schema: `{type, id, agent, to, target, msg, tags, ts}`. Lifecycle:
`send_start` → `sidechat_autoprovisioned` → `verified` → `sent`.
- Dispatch: `dm.py send --agent <dispatcher> --to <worker> --target sw-<id>-s<slot>`
with body `[JOB sw-<id>/<slot>] Task for swarm slot N:\n<task>\n\nReply with [RESULT sw-...]`
- Target auto-provisions a sidechat; 23 slot sidechats are registered in
`job-sidechats.json` with per-slot `agent` attribution (e.g. `agent: "dev"`).
- Harvester (`response-harvester.py:773`) bridges sidechat
`[RESULT <swarm_id>/<slot>]` markers into `box swarm-report` — but the
sibling's `reporter.py` calls `swarm-report` directly (no DM round-trip).
## Recommendation: (a) direct sidechat post, main loop as visibility layer
**Result recording (unchanged):** worker → `box swarm-report` directly
(sibling's reporter.py). Reliable, no harvester-scan dependency.
**Main-loop bridge (new, zero code changes):** after recording the result,
the worker posts a short human-readable completion note to the slot's own
sidechat via `dm.py send` — deliberately *without* a `[JOB ` tag so the
main loop's `check_sidechats()` does not filter it:
```
[SWARM-DONE sw-20261005-151307-a222/0] OK in 12.4s: <first ~200 chars of output>
```
Why this works with no modifications:
1. Slot sidechats are already registered in `job-sidechats.json` with agent
attribution → `get_monitored_sidechats()` includes them for that agent.
2. The note contains no `[JOB ` → not skipped by the message filter.
3. Next main-loop tick picks it up → digest → agent's prompting sidechat →
the agent *sees* the completion in the same surface as everything else.
This is "answering via the main loop": completions surface through the
existing digest machinery instead of a parallel notification channel.
**Why not (b) queue via `send_prompt()`:** `send_prompt()` is built for
main-chat digests → prompting sidechats, with `[JOB ml-...]` followup
semantics. Swarm results are not digests; routing them through would add
timer latency, misuse the ml- followup namespace, and conflate two concerns.
**Why not (c) worker as main-loop participant:** the main loop monitors
muse.ai *agent accounts* (`muse_hybrid.get_history(agent, ...)`). The worker
is a bl daemon with no agent account. Participation-by-proxy (writing to
monitored sidechats) achieves the same visibility with no identity hacks.
## Watchdog (no change needed, documented here)
`check_sidechats()` also gives stall detection for free: a claimed slot
whose sidechat shows no `[SWARM-DONE]`/`[RESULT]` within the slot timeout
is visible as inactivity in the digest. A future enhancement (not this
bridge) could add an explicit stall-threshold alert.
## Interface for the worker daemon
After `reporter.post_result(...)` succeeds, call:
`notify_via_mainloop(swarm_id, slot_index, summary_text)` (prototype:
`bin/swarm_worker/mainloop_notify.py`). It resolves
`sw-<swarm_id>-s<slot_index>` → sidechat UUID via `dm.resolve_sidechat_target`,
then `dm.py send` with the `[SWARM-DONE ...]` note. Dry-run mode posts nothing.
## Files
- This doc: `bin/swarm_worker/mainloop-bridge.md`
- Prototype: `bin/swarm_worker/mainloop_notify.py` (dry-run safe, never auto-wired)