feat(swarm): launch live swarm worker daemon under tmux supervisor
This commit is contained in:
@@ -0,0 +1,95 @@
|
||||
# Swarm Worker ↔ Main Loop Bridge — Design Doc
|
||||
**Bridge Builder 5/5 | 2026-10-05 | Read-only design (no changes to self_main_loop.py)**
|
||||
|
||||
## Context
|
||||
|
||||
Sibling builders produced a tmux-hosted worker daemon for bl:
|
||||
- `executor.py` — sandboxed task execution (shell vs prose heuristic, refuse-list)
|
||||
- `poller.py` — slot claiming
|
||||
- `reporter.py` — result recording via `box swarm-report <id> <slot>` (direct, reliable)
|
||||
|
||||
The user's directive: *"we should be answering our own main chat via our main loop."*
|
||||
This doc designs how swarm work becomes visible through `self_main_loop.py`
|
||||
without modifying it.
|
||||
|
||||
## How the main loop works (verified against live code)
|
||||
|
||||
`bin/self_main_loop.py` (995 lines) is a timer-driven read-side loop:
|
||||
1. Reads each watched agent's muse.ai **main chat** + registered **sidechats**
|
||||
(via `muse_hybrid` gateway, `box-chat.py` fallback).
|
||||
2. On new activity, composes a digest and posts it to the agent's **prompting
|
||||
sidechat** via `dm.py send` (`send_prompt()`).
|
||||
3. Actionable digests (containing `(?)` or `(!)`) get `[JOB ml-<agent>-<ts>]`
|
||||
+ contract footer → `--expect-reply` → followup record. Informational
|
||||
digests go via fast gateway, no followup.
|
||||
4. Response verbs `[ACK|CLAIM|RESULT|DECLINE|NO-ACTION <id>]` are matched by
|
||||
`response-harvester.py` to resolve followups.
|
||||
|
||||
Key detail — `check_sidechats()` **skips individual messages** containing
|
||||
`"[JOB "`, `"Heartbeat check"`, or `"[from:super]"`. It does NOT skip entire
|
||||
sidechats. A `[RESULT ...]` or `[SWARM-DONE ...]` post is *not* filtered.
|
||||
|
||||
## How swarm slot DMs arrive (verified against dm-log.jsonl)
|
||||
|
||||
Schema: `{type, id, agent, to, target, msg, tags, ts}`. Lifecycle:
|
||||
`send_start` → `sidechat_autoprovisioned` → `verified` → `sent`.
|
||||
|
||||
- Dispatch: `dm.py send --agent <dispatcher> --to <worker> --target sw-<id>-s<slot>`
|
||||
with body `[JOB sw-<id>/<slot>] Task for swarm slot N:\n<task>\n\nReply with [RESULT sw-...]`
|
||||
- Target auto-provisions a sidechat; 23 slot sidechats are registered in
|
||||
`job-sidechats.json` with per-slot `agent` attribution (e.g. `agent: "dev"`).
|
||||
- Harvester (`response-harvester.py:773`) bridges sidechat
|
||||
`[RESULT <swarm_id>/<slot>]` markers into `box swarm-report` — but the
|
||||
sibling's `reporter.py` calls `swarm-report` directly (no DM round-trip).
|
||||
|
||||
## Recommendation: (a) direct sidechat post, main loop as visibility layer
|
||||
|
||||
**Result recording (unchanged):** worker → `box swarm-report` directly
|
||||
(sibling's reporter.py). Reliable, no harvester-scan dependency.
|
||||
|
||||
**Main-loop bridge (new, zero code changes):** after recording the result,
|
||||
the worker posts a short human-readable completion note to the slot's own
|
||||
sidechat via `dm.py send` — deliberately *without* a `[JOB ` tag so the
|
||||
main loop's `check_sidechats()` does not filter it:
|
||||
|
||||
```
|
||||
[SWARM-DONE sw-20261005-151307-a222/0] OK in 12.4s: <first ~200 chars of output>
|
||||
```
|
||||
|
||||
Why this works with no modifications:
|
||||
1. Slot sidechats are already registered in `job-sidechats.json` with agent
|
||||
attribution → `get_monitored_sidechats()` includes them for that agent.
|
||||
2. The note contains no `[JOB ` → not skipped by the message filter.
|
||||
3. Next main-loop tick picks it up → digest → agent's prompting sidechat →
|
||||
the agent *sees* the completion in the same surface as everything else.
|
||||
This is "answering via the main loop": completions surface through the
|
||||
existing digest machinery instead of a parallel notification channel.
|
||||
|
||||
**Why not (b) queue via `send_prompt()`:** `send_prompt()` is built for
|
||||
main-chat digests → prompting sidechats, with `[JOB ml-...]` followup
|
||||
semantics. Swarm results are not digests; routing them through would add
|
||||
timer latency, misuse the ml- followup namespace, and conflate two concerns.
|
||||
|
||||
**Why not (c) worker as main-loop participant:** the main loop monitors
|
||||
muse.ai *agent accounts* (`muse_hybrid.get_history(agent, ...)`). The worker
|
||||
is a bl daemon with no agent account. Participation-by-proxy (writing to
|
||||
monitored sidechats) achieves the same visibility with no identity hacks.
|
||||
|
||||
## Watchdog (no change needed, documented here)
|
||||
|
||||
`check_sidechats()` also gives stall detection for free: a claimed slot
|
||||
whose sidechat shows no `[SWARM-DONE]`/`[RESULT]` within the slot timeout
|
||||
is visible as inactivity in the digest. A future enhancement (not this
|
||||
bridge) could add an explicit stall-threshold alert.
|
||||
|
||||
## Interface for the worker daemon
|
||||
|
||||
After `reporter.post_result(...)` succeeds, call:
|
||||
`notify_via_mainloop(swarm_id, slot_index, summary_text)` (prototype:
|
||||
`bin/swarm_worker/mainloop_notify.py`). It resolves
|
||||
`sw-<swarm_id>-s<slot_index>` → sidechat UUID via `dm.resolve_sidechat_target`,
|
||||
then `dm.py send` with the `[SWARM-DONE ...]` note. Dry-run mode posts nothing.
|
||||
|
||||
## Files
|
||||
- This doc: `bin/swarm_worker/mainloop-bridge.md`
|
||||
- Prototype: `bin/swarm_worker/mainloop_notify.py` (dry-run safe, never auto-wired)
|
||||
Reference in New Issue
Block a user