97 lines
4.3 KiB
Markdown
97 lines
4.3 KiB
Markdown
# Supervision Scope Contract
|
||
|
||
Status: **Draft** — decisions below are unsettled until marked otherwise.
|
||
Only explicit user acceptance moves this document (or any decision) to Final.
|
||
|
||
## Goal
|
||
|
||
Every fleet node stays alive and truthfully reported: browsers supervised,
|
||
relays supervised, dead nodes recovered or loudly paged, and `box` status
|
||
honest from any shell (including PID/net-blind sandboxed shells).
|
||
|
||
## Non-goals (proposed)
|
||
|
||
- Agent lifecycle/onboarding stages (provision, auth, OTP, invite redeem).
|
||
- Work completion (job dispatch, followups, harvester, completion auditor).
|
||
- Loop-health scoring and drive repair.
|
||
|
||
## Supervisors (observed, all installed 2026-10-06)
|
||
|
||
| Supervisor | Cadence | Coverage | Decides |
|
||
|---|---|---|---|
|
||
| chromebox-watchdog@\<node\>.timer ×6 | 2 min | all registry nodes (def/dev timers installed 18:32Z) | browser+egress per node; tunnel restart, chrome relaunch |
|
||
| cdp-relay-watchdog.timer | 5 min | registry-driven (`watched_nodes()`) | relay veth IP + connectivity; relay restart |
|
||
| agent-health.timer (user) | 5 min | registry-driven | API check per node; kill+restart with 2-strike rule + circuit breaker (3 futile → open 30 min) |
|
||
| ensure-node-supervision.sh | on node-up / `--all` | new + drifted nodes | NODES.md row + chromebox timer install |
|
||
| host_evidence fallback | on `box` read | registry (relay) + installed timers (browser) | effective status when live probes are blind |
|
||
|
||
## Decisions
|
||
|
||
(D1..D7 below — all UNRESOLVED unless marked.)
|
||
|
||
### D1. Contract boundary: which supervisors are in scope — SETTLED (recommended accepted)
|
||
|
||
IN: chromebox-watchdog ×6, cdp-relay-watchdog, agent-health + circuit
|
||
breaker, ensure-node-supervision feed, host_evidence fallback.
|
||
OUT: onboarding pipeline lifecycle, completion auditor, loop-health
|
||
(each keeps its own owner and interviews separately).
|
||
|
||
### D2. Kill-path precedence (chromebox-watchdog vs agent-health) — SETTLED (recommended accepted)
|
||
|
||
Policy: each supervisor owns its signal — chromebox-watchdog owns
|
||
CDP/egress/process failure, agent-health owns API-layer failure.
|
||
Browser restart stays the shared actuator, and both keep their <2-min
|
||
fresh-browser guards. No code change; this paragraph is the guard
|
||
against future edits dropping either side.
|
||
|
||
### D3. Egress-down fall-through (relaunch chrome after failed tunnel restart?) — SETTLED (recommended accepted)
|
||
|
||
Policy: stop after a failed tunnel restart. Log the egress failure
|
||
loudly and page it as a tunnel fault, never as a browser fault; skip
|
||
the chrome relaunch. The next 2-min run retries the tunnel. (Requires
|
||
a chromebox-watchdog.sh change — implementation needs a separate
|
||
explicit request.)
|
||
|
||
### D4. Circuit-breaker thresholds (3 futile / 30 min cooldown) — SETTLED (recommended accepted)
|
||
|
||
Keep 3 strikes / 1800s cooldown as initial values, with a review
|
||
trigger: revisit if one node trips the circuit more than twice in a
|
||
week (threshold too touchy) or a stuck node sits out a full cooldown
|
||
under operator eyes (cooldown too long).
|
||
|
||
### D5. Coverage source of truth — SETTLED (recommended accepted)
|
||
|
||
Registry-only: every active NODES.md row is supervised, period. The
|
||
ensure script guarantees the timer follows the row, and a missing
|
||
timer degrades to an honest `unknown` (no journal runs) rather than a
|
||
wrong verdict. No second inventory.
|
||
|
||
### D6. Concurrent-edit protocol for shared supervision files — SETTLED (recommended accepted)
|
||
|
||
Protocol: claim-before-edit (announce shared-file claims in chat),
|
||
anchored-edits-only on shared files (never full-file rewrites, so
|
||
concurrent work interleaves), re-read before staging, and prompt
|
||
mine-only commits so landed work survives the next clobber.
|
||
|
||
### D7. Done means — UNRESOLVED
|
||
|
||
Proposed checklist: timers on all 6 firing silent; relay/agent-health
|
||
loops registry-driven with tests; ensure hook live; this doc Final.
|
||
|
||
## Risks
|
||
|
||
- Egress-down pages read as browser failures (D3).
|
||
- Uncommitted supervisor work can be clobbered by a concurrent editor (D6).
|
||
- c9143a5 contains unattributed peer hunks (host_evidence `_covered_nodes`,
|
||
fleet-status test updates) — needs an amend-or-leave decision.
|
||
|
||
## Validation
|
||
|
||
- `box fleet status` truthful from blind shells (live-verified 6/6 ACTIVE).
|
||
- Focused suites green (supervision, fleet, agent-health, watchdog coverage).
|
||
- Timer firing proven via journal, not config presence.
|
||
|
||
## Unresolved items
|
||
|
||
D7 unresolved. D1–D6 settled.
|