- bin/muse_choice_watcher.py + systemd/muse-choices-reconcile.*: automatic choice answering and timer reconciliation - bin/digest.py: fleet log and health summarization - docs/BOX-*-HTTPS.md: comprehensive HTTPS execution contracts and API documentation - docs/MUSE-CHOICES-POLICY.md & docs/SUPERVISION-SPEC.md: autonomous execution specs - tests/test_*.py: unit test suites for HTTPS API, choice watcher, fleet heal, and swarm pruning
3.4 KiB
Supervision Scope Contract
Status: Draft — decisions below are unsettled until marked otherwise. Only explicit user acceptance moves this document (or any decision) to Final.
Goal
Every fleet node stays alive and truthfully reported: browsers supervised,
relays supervised, dead nodes recovered or loudly paged, and box status
honest from any shell (including PID/net-blind sandboxed shells).
Non-goals (proposed)
- Agent lifecycle/onboarding stages (provision, auth, OTP, invite redeem).
- Work completion (job dispatch, followups, harvester, completion auditor).
- Loop-health scoring and drive repair.
Supervisors (observed, all installed 2026-10-06)
| Supervisor | Cadence | Coverage | Decides |
|---|---|---|---|
| chromebox-watchdog@<node>.timer ×6 | 2 min | all registry nodes (def/dev timers installed 18:32Z) | browser+egress per node; tunnel restart, chrome relaunch |
| cdp-relay-watchdog.timer | 5 min | registry-driven (watched_nodes()) |
relay veth IP + connectivity; relay restart |
| agent-health.timer (user) | 5 min | registry-driven | API check per node; kill+restart with 2-strike rule + circuit breaker (3 futile → open 30 min) |
| ensure-node-supervision.sh | on node-up / --all |
new + drifted nodes | NODES.md row + chromebox timer install |
| host_evidence fallback | on box read |
registry (relay) + installed timers (browser) | effective status when live probes are blind |
Decisions
(D1..D7 below — all UNRESOLVED unless marked.)
D1. Contract boundary: which supervisors are in scope — SETTLED (recommended accepted)
IN: chromebox-watchdog ×6, cdp-relay-watchdog, agent-health + circuit breaker, ensure-node-supervision feed, host_evidence fallback. OUT: onboarding pipeline lifecycle, completion auditor, loop-health (each keeps its own owner and interviews separately).
D2. Kill-path precedence (chromebox-watchdog vs agent-health) — UNRESOLVED
Both can kill a browser today; only time guards (<2 min) de-conflict them.
D3. Egress-down fall-through (relaunch chrome after failed tunnel restart?) — UNRESOLVED
Observed 19:12Z: tunnel restart failed, watchdog relaunched chrome 3× anyway (one FAILED page). Browser was never the problem.
D4. Circuit-breaker thresholds (3 futile / 30 min cooldown) — UNRESOLVED
Current values unvalidated against real recurrence intervals.
D5. Coverage source of truth — UNRESOLVED
Registry-only vs registry+installed-timers for browser verdicts.
D6. Concurrent-edit protocol for shared supervision files — UNRESOLVED
Two agents editing super-cli.py / watchdogs / runbook; one clobber
(18:03Z) and one unattributed commit (c9143a5) already occurred.
D7. Done means — UNRESOLVED
Proposed checklist: timers on all 6 firing silent; relay/agent-health loops registry-driven with tests; ensure hook live; this doc Final.
Risks
- Egress-down pages read as browser failures (D3).
- Uncommitted supervisor work can be clobbered by a concurrent editor (D6).
c9143a5contains unattributed peer hunks (host_evidence_covered_nodes, fleet-status test updates) — needs an amend-or-leave decision.
Validation
box fleet statustruthful from blind shells (live-verified 6/6 ACTIVE).- Focused suites green (supervision, fleet, agent-health, watchdog coverage).
- Timer firing proven via journal, not config presence.
Unresolved items
D2–D7 unresolved. D1 settled.