feat: completion-enforcement loop (fallback, proof, emit-model, auditor)

Close the loop so dispatched work actually completes on bl:

- on_no_result fallback in followup-sweeper (op + job forms via
  exec-constrained registry / job-dispatch), seeded on the three
  autonomy-pulse jobs; fallback_due() dedupes the gravity path
- gravity.py: add __main__ entry (loop-remediator.timer was a no-op),
  300s re-arm budget, fallback firing + stamp/skip logic
- harvester: proof-of-result followups (result_has_evidence),
  acted-variant NACK, emit-model tool-hint wording
- envelope: RESPONSE RULE states the emit model (agents EMIT
  directives verbatim; runtime executes; works from bare containers)
- completion-audit.py + systemd 15-min timer: per-family funnel,
  swarm drain, followup backlog; digest DM when degraded, 6h heartbeat
- tests/test_completion.py (29 tests), JOB-SPEC.md docs

Tests: 67/67 focused green (completion + tool_calls).
This commit is contained in:
Muse Sidechat
2026-10-06 08:20:14 +00:00
parent b7e45010c3
commit a9f014f9fa
12 changed files with 1173 additions and 16 deletions
+80
View File
@@ -225,6 +225,86 @@ spam if a job misfires in a loop.
- DMs are logged (dm-log.jsonl) for audit
- Side chats are per-job, not shared across trust boundaries
## Completion Enforcement
Sending a job DM does not complete work. Four mechanisms close the loop on bl:
### 1. `on_no_result` fallback (server-side guarantee)
A job may declare a fallback effect that runs when its followup reaches
terminal expiry with no agent result (`bin/followup-sweeper.py`):
```json
"on_no_result": {"op": "swarm.spawn", "args": {"count": 2, "task": "...", "label": "..."}}
"on_no_result": {"job": "<another-job-name>"}
```
- `op` form: validated + built through the `exec-constrained` op registry
(same validators the daemon uses), then executed as a subprocess.
- `job` form: dispatches the named job via `job-dispatch.py` with
`CHAIN_PREV_*` timeout context.
- Outcomes log as `fallback_executed` / `fallback_failed` in `job-log.jsonl`
and stamp `rec["fallback"]` on the followup record. Failures never break
the sweep. Jobs without the key behave exactly as before.
- Firing paths (either; never both): the sweeper terminal branch
(`nudges_sent >= nudges_allowed`, expired) and `gravity.py:remediate_breaks`
(same terminal condition, runs on `loop-remediator.timer` ~every 15m).
`fallback_due()` dedupes: fires once per record, retries a failed attempt
after 1h. Gravity also revives the local sweep loop (it previously had no
`__main__`, so the timer was a no-op) with a 300s re-arm budget (was 10s,
which strangled every sweep mid-first-send — median nudge send is ~10s).
- Split-brain note: the VM board sweeper (`box-request-sweeper.timer`) sends
the live `[NUDGE <id>]` DMs from remote `dm_followup` requests; the local
sweeper sends `[nudge N/M]` from `followups.json`. Both fire per deadline
until a reply resolves both sides. Accept the duplicate nudge as the cost
of a guaranteed local path; cross-system dedup is future work.
- Seeded on: `autonomy-pulse-646`, `autonomy-pulse-pip`, `autonomy-pulse-opm`
(each spawns 2 standing-work subagents if the agent naps through the pulse).
- Race note: set the job's followup timeout longer than the expected work
time, or a slow-but-working agent can double-fire alongside the fallback.
### 2. NACK for directive-less replies (existing, sharpened)
`response-harvester.py:maybe_nudge_untagged_sidechat` already rejects
conversational replies in job-backed threads (1-turn strict nudge, then a
5-minute escalation timer to opm; tracked in
`conversation-nudge-tracker.json`). Two refinements:
- Acted-variant: when the agent emitted directives but never closed with
`[RESULT]`, the nudge acknowledges the action and demands the close
instead of crying "commentary rejected".
- Emit-model wording: envelope (`prompt_envelope._response_rule`), tool
hint, and strict nudge now state plainly that agents EMIT `[TOOL]` /
`[DM]` lines verbatim and the Box runtime on bl executes them — this
works from containers with no box CLI or tmux socket. (Root cause of the
2026-10-06 zero-`tool_exec` stretch: agents believed they had to execute
tools locally and declined for lack of a "container equivalent".)
### 3. Proof-of-result followups
A success `[RESULT]` with no checkable artifact (swarm/timer IDs, UUIDs,
paths, `n/m` completion counts — see `result_has_evidence`) triggers a
one-shot `followup.create` (+30m, same thread) asking for the evidence, and
logs `proof_requested`. One per (thread, job). Failures and declines skip
proof (they already chain via `on_failure`).
### 4. Completion auditor (15-minute timer)
`bin/completion-audit.py` via `systemd/completion-audit.timer`
(`OnCalendar=*:4/15`, installed to `/etc/systemd/system`):
- Computes the funnel per job family over `--hours` (default 24) from
`job-log.jsonl`: sent / dispatched / `tool_exec` / results, plus
`fallback_*`, `proof_requested`, and `job_failed` counts.
- Adds point-in-time swarm slot drain (flags running slots idle >90m) and
followup backlog (pending / overdue / escalated).
- Writes `logs/completion-audit-<ts>.json` + `logs/completion-audit-latest.json`.
- Posts the digest to `opm` / `heartbeat` only when DEGRADED
(silent families, failures, tool errors, stale slots, overdue/escalated
followups); when green, posts a heartbeat at most every 6h
(`logs/completion-audit-state.json`). Silent-when-healthy otherwise.
- Read-only except the digest DM and its own log/state files.
## Future Expansions
1. **Conditional jobs**: Run Job B only if Job A succeeds with specific output