feat: completion-enforcement loop (fallback, proof, emit-model, auditor)
Close the loop so dispatched work actually completes on bl: - on_no_result fallback in followup-sweeper (op + job forms via exec-constrained registry / job-dispatch), seeded on the three autonomy-pulse jobs; fallback_due() dedupes the gravity path - gravity.py: add __main__ entry (loop-remediator.timer was a no-op), 300s re-arm budget, fallback firing + stamp/skip logic - harvester: proof-of-result followups (result_has_evidence), acted-variant NACK, emit-model tool-hint wording - envelope: RESPONSE RULE states the emit model (agents EMIT directives verbatim; runtime executes; works from bare containers) - completion-audit.py + systemd 15-min timer: per-family funnel, swarm drain, followup backlog; digest DM when degraded, 6h heartbeat - tests/test_completion.py (29 tests), JOB-SPEC.md docs Tests: 67/67 focused green (completion + tool_calls).
This commit is contained in:
@@ -225,6 +225,86 @@ spam if a job misfires in a loop.
|
||||
- DMs are logged (dm-log.jsonl) for audit
|
||||
- Side chats are per-job, not shared across trust boundaries
|
||||
|
||||
## Completion Enforcement
|
||||
|
||||
Sending a job DM does not complete work. Four mechanisms close the loop on bl:
|
||||
|
||||
### 1. `on_no_result` fallback (server-side guarantee)
|
||||
|
||||
A job may declare a fallback effect that runs when its followup reaches
|
||||
terminal expiry with no agent result (`bin/followup-sweeper.py`):
|
||||
|
||||
```json
|
||||
"on_no_result": {"op": "swarm.spawn", "args": {"count": 2, "task": "...", "label": "..."}}
|
||||
"on_no_result": {"job": "<another-job-name>"}
|
||||
```
|
||||
|
||||
- `op` form: validated + built through the `exec-constrained` op registry
|
||||
(same validators the daemon uses), then executed as a subprocess.
|
||||
- `job` form: dispatches the named job via `job-dispatch.py` with
|
||||
`CHAIN_PREV_*` timeout context.
|
||||
- Outcomes log as `fallback_executed` / `fallback_failed` in `job-log.jsonl`
|
||||
and stamp `rec["fallback"]` on the followup record. Failures never break
|
||||
the sweep. Jobs without the key behave exactly as before.
|
||||
- Firing paths (either; never both): the sweeper terminal branch
|
||||
(`nudges_sent >= nudges_allowed`, expired) and `gravity.py:remediate_breaks`
|
||||
(same terminal condition, runs on `loop-remediator.timer` ~every 15m).
|
||||
`fallback_due()` dedupes: fires once per record, retries a failed attempt
|
||||
after 1h. Gravity also revives the local sweep loop (it previously had no
|
||||
`__main__`, so the timer was a no-op) with a 300s re-arm budget (was 10s,
|
||||
which strangled every sweep mid-first-send — median nudge send is ~10s).
|
||||
- Split-brain note: the VM board sweeper (`box-request-sweeper.timer`) sends
|
||||
the live `[NUDGE <id>]` DMs from remote `dm_followup` requests; the local
|
||||
sweeper sends `[nudge N/M]` from `followups.json`. Both fire per deadline
|
||||
until a reply resolves both sides. Accept the duplicate nudge as the cost
|
||||
of a guaranteed local path; cross-system dedup is future work.
|
||||
- Seeded on: `autonomy-pulse-646`, `autonomy-pulse-pip`, `autonomy-pulse-opm`
|
||||
(each spawns 2 standing-work subagents if the agent naps through the pulse).
|
||||
- Race note: set the job's followup timeout longer than the expected work
|
||||
time, or a slow-but-working agent can double-fire alongside the fallback.
|
||||
|
||||
### 2. NACK for directive-less replies (existing, sharpened)
|
||||
|
||||
`response-harvester.py:maybe_nudge_untagged_sidechat` already rejects
|
||||
conversational replies in job-backed threads (1-turn strict nudge, then a
|
||||
5-minute escalation timer to opm; tracked in
|
||||
`conversation-nudge-tracker.json`). Two refinements:
|
||||
|
||||
- Acted-variant: when the agent emitted directives but never closed with
|
||||
`[RESULT]`, the nudge acknowledges the action and demands the close
|
||||
instead of crying "commentary rejected".
|
||||
- Emit-model wording: envelope (`prompt_envelope._response_rule`), tool
|
||||
hint, and strict nudge now state plainly that agents EMIT `[TOOL]` /
|
||||
`[DM]` lines verbatim and the Box runtime on bl executes them — this
|
||||
works from containers with no box CLI or tmux socket. (Root cause of the
|
||||
2026-10-06 zero-`tool_exec` stretch: agents believed they had to execute
|
||||
tools locally and declined for lack of a "container equivalent".)
|
||||
|
||||
### 3. Proof-of-result followups
|
||||
|
||||
A success `[RESULT]` with no checkable artifact (swarm/timer IDs, UUIDs,
|
||||
paths, `n/m` completion counts — see `result_has_evidence`) triggers a
|
||||
one-shot `followup.create` (+30m, same thread) asking for the evidence, and
|
||||
logs `proof_requested`. One per (thread, job). Failures and declines skip
|
||||
proof (they already chain via `on_failure`).
|
||||
|
||||
### 4. Completion auditor (15-minute timer)
|
||||
|
||||
`bin/completion-audit.py` via `systemd/completion-audit.timer`
|
||||
(`OnCalendar=*:4/15`, installed to `/etc/systemd/system`):
|
||||
|
||||
- Computes the funnel per job family over `--hours` (default 24) from
|
||||
`job-log.jsonl`: sent / dispatched / `tool_exec` / results, plus
|
||||
`fallback_*`, `proof_requested`, and `job_failed` counts.
|
||||
- Adds point-in-time swarm slot drain (flags running slots idle >90m) and
|
||||
followup backlog (pending / overdue / escalated).
|
||||
- Writes `logs/completion-audit-<ts>.json` + `logs/completion-audit-latest.json`.
|
||||
- Posts the digest to `opm` / `heartbeat` only when DEGRADED
|
||||
(silent families, failures, tool errors, stale slots, overdue/escalated
|
||||
followups); when green, posts a heartbeat at most every 6h
|
||||
(`logs/completion-audit-state.json`). Silent-when-healthy otherwise.
|
||||
- Read-only except the digest DM and its own log/state files.
|
||||
|
||||
## Future Expansions
|
||||
|
||||
1. **Conditional jobs**: Run Job B only if Job A succeeds with specific output
|
||||
|
||||
Reference in New Issue
Block a user