E1-E6: supervised 3-observation proof bar per kind, independent flips, automatic on proof, single-miss rollback, fleet-wide, grill questions stay coordinator-gated. Scope accepted verbatim in-record.
8.5 KiB
title, status, coordinator, scope, accepted_at, accepted_quote, gate, signoff_targets
| title | status | coordinator | scope | accepted_at | accepted_quote | gate | signoff_targets | |||
|---|---|---|---|---|---|---|---|---|---|---|
| Muse-Choices Deny/Escalate & Agy Hold-All Exit Policy | signed-off | operator-main | muse-choices-policy | 2026-10-07T05:55:03Z | ACCEPT | coordinator |
|
Muse-Choices Deny/Escalate Policy — DECISION RECORD (Final)
Topic: add deny/escalate decisions to the muse-choices auto-approve daemon
(bin/muse_choice_watcher.py), which today only approves (top choice per
prompt kind). Interviewed 2026-10-06/07 per grill contract; accepted
verbatim below, which flipped this record from Draft to Final.
Standing constraints (settled by user)
- All prompts must resolve: no stuck states are acceptable in any outcome.
- The full-auto top-choice flow must always exist as a path.
- Model review of choices is a FUTURE layer. Deferred out of this interview.
Settled decisions
- D0 (helper form): a checked-in repo rules file informs decisions. Source: user's structured answerquared 2026-10-06 ("Repo rules file (Recommended)" for "which helper should inform approve/deny/hold decisions"). Rationale recorded at selection time: deterministic, versioned, sub-second, unit-testable; no agent round-trip latency.
Settled during interview
-
D0b (approve path needs no helper): straightforward prompts resolve locally with top-choice keys (the five matcher kinds, already implemented and live). Helpers (D0 rules file) govern deny/hold judgments only. Source: user direction 2026-10-06 ("the watcher itself should be able to input 1"; "always have the flow for full auto just top choice").
-
D1 (rule match dimensions): command pattern first, plus kind and text pattern. Source: user selected option 1, 2026-10-06. Rationale: the observed risk lives in the
$ commandof approval dialogs; kind/text add precision around it. -
D2 (deny mechanics): deny exists only for permission kinds --
muse-approvaldialogs receive2+ Enter,y/nprompts receiven+ Enter. Question kinds (interview, letter, numbered) always resolve top-choice and are never denied. Source: user selected option 1, 2026-10-06. Rationale: deny is only meaningful where a permission is refused; questions stay total. -
D3 (hold mechanics): hold leaves the dialog untouched, suppresses auto-answer, raises a HELD entry in
box muse-choices statusplus an audit record; the operator resolves via a box command, otherwise a SHORT window expires back to top-choice approve. Source: user selected option 1 with "short window", 2026-10-06. Exact duration proposed below (2 minutes, tunable); accepted or amended with the scope text in D5. -
D4 (unmatched default): approve top-choice, exactly today's behavior. Source: user selected option 1, 2026-10-06. Rationale: follows from the standing constraints; rules carve out only deny/hold exceptions, so an empty rules file changes nothing.
Open questions (unresolved)
None. All interview questions resolved and the scope accepted.
Scope contract (ACCEPTED)
Artifact boundary, IN:
docs/MUSE-CHOICES-POLICY.md: this record (Draft -> Final on acceptance).- New
muse-choices-rules.jsonat repo root (besidekeepalive-config.json): the checked-in deny/hold rules. bin/muse_choice_watcher.py: rule evaluation, deny/hold paths, HELD state with short-window expiry, resolve plumbing.bin/super-cli.py:box muse-choices resolvecommand + HELD display in status.tests/test_muse_choice_watcher.py: rule eval, per-kind deny keys, hold/suppress/expiry, resolve flow.
Artifact boundary, OUT (rejected or deferred, each needs its own interview to re-enter):
- Model review of choices (deferred future stage).
- Peer-agent consultation (rejected in D0).
- New matcher shapes (matcher suite's lane).
box runtimework (adjacent lane, untouched).- Timer cadence / daemon supervision changes.
Done means (all observable):
- This record marked Final with the acceptance quoted.
- Rules file loads; empty rules == today's behavior exactly.
- Deny sends
2+Enter /n+Enter per D2: unit tests + one live scratch proof per permission kind. - Hold suppresses + shows HELD + resolves via box + expires to approve: unit tests + one live scratch proof of hold and one of expiry.
- Audit records for deny/hold/resolve/expire in
box-ctl.jsonl. - Full suite green; fleet reloaded; desired state left as found.
Acceptance (quoted verbatim, chat, 2026-10-07T00:19:57Z): "ACCEPT". Accepted as written, including the 2-minute tunable hold window. Per the grill scope contract, later work outside the IN boundary needs explicit owner approval or its own follow-up interview; "go" authorizes only this boundary. No owning issue exists in this workflow, so this record is the lane-coordination evidence.
Non-goals (accepted with the scope)
- Model-based review of choices (deferred future layer).
- Peer-agent consultation over sidechat (rejected in favor of D0).
- Changes to approval matching shapes (covered by the matcher test suite).
agy hold-all exit criteria — DECISION RECORD (Final)
Interview opened 2026-10-07. Topic selected by user (option 1, chat): what proof flips agy panes from hold-all to auto-answering. The D0–D5 record above is Final and untouched by this interview. Numbering continues as E1+ to avoid collision.
Settled during interview
- E1 (proof bar): supervised live proof per kind, fixed count. Source: user selected option 1, 2026-10-07. Rationale: the risk is behavioral (how agy's renderer treats digit keys) — only live observation on real agy prompts proves it; synthetic tests cannot. A correct observation means a resolve-approve (or equivalent human-gated send of the kind's key) produces the intended selection with no stray keys. The count itself is E1b.
- E1b (count): three consecutive correct observations per kind, uniform across kinds; a miss resets the count. Source: user selected option 1, 2026-10-07. Rationale: one success could be luck (cursor already placed); three in a row across separate prompts proves the key binding, not the moment.
- E2 (rollout grain): independent per-kind flip — each kind exits hold-all the moment its 3 observations complete. Source: user selected option 1, 2026-10-07. Rationale: matches the uniform bar; rarely-seen kinds do not block proven ones.
- E3 (flip authority): automatic on completed proof — the 3 observations are themselves human-gated (each a resolve-approve), so completion is the approval. The flip applies automatically and is recorded in the audit log plus a dated note in this doc. Source: user selected option 1, 2026-10-07. Rationale: no extra ceremony beyond the supervision already done; the audit trail shows all 3 proofs.
- E4 (rollback): single miss re-holds the kind — any post-flip wrong-send (wrong selection or stray keys), however observed, immediately re-holds that kind and wipes its proof count; re-flip needs 3 fresh observations. Source: user selected option 1, 2026-10-07. Rationale: fail-closed and symmetric with the bar.
- E5 (applicability): fleet-wide — proof on any agy pane flips the kind on all agy panes. Source: user selected option 1, 2026-10-07. Rationale: the key binding is a property of the agy binary, identical on every pane.
- E6 (grill exclusion): numbered-plain stays coordinator-gated on agy; the E-rules never flip it. Source: user selected option 1, 2026-10-07. Rationale: the E-rules prove the technical question (keys work); the grill hold answers the governance question (interviews need sign-off). Proof cannot lift a governance gate.
Open questions (unresolved)
None. All interview questions resolved and the scope accepted.
Scope contract (ACCEPTED)
- IN: this record section only (agy exit criteria, now Final).
- OUT: the flip mechanism (code), tests, daemon restarts, any implementation. Accepting this record never approves those stages; each returns for its own go-ahead.
- Done means: every E-question settled below, section marked Final, user acceptance quoted verbatim with channel and time.
- "Go" / "do it all" authorize only the IN boundary. No owning issue exists in this workflow, so this record is the lane-coordination evidence.
Acceptance (quoted verbatim, chat, 2026-10-07T05:55:03Z): "ACCEPT". Accepted as written. Per the grill scope contract, later work outside the IN boundary (the flip mechanism, tests, restarts) needs explicit owner approval or its own follow-up interview; "go" authorizes only this boundary.