Files
box/docs/SUPERVISION-SPEC.md
T

4.3 KiB
Raw Blame History

Supervision Scope Contract

Status: Draft — decisions below are unsettled until marked otherwise. Only explicit user acceptance moves this document (or any decision) to Final.

Goal

Every fleet node stays alive and truthfully reported: browsers supervised, relays supervised, dead nodes recovered or loudly paged, and box status honest from any shell (including PID/net-blind sandboxed shells).

Non-goals (proposed)

  • Agent lifecycle/onboarding stages (provision, auth, OTP, invite redeem).
  • Work completion (job dispatch, followups, harvester, completion auditor).
  • Loop-health scoring and drive repair.

Supervisors (observed, all installed 2026-10-06)

Supervisor Cadence Coverage Decides
chromebox-watchdog@<node>.timer ×6 2 min all registry nodes (def/dev timers installed 18:32Z) browser+egress per node; tunnel restart, chrome relaunch
cdp-relay-watchdog.timer 5 min registry-driven (watched_nodes()) relay veth IP + connectivity; relay restart
agent-health.timer (user) 5 min registry-driven API check per node; kill+restart with 2-strike rule + circuit breaker (3 futile → open 30 min)
ensure-node-supervision.sh on node-up / --all new + drifted nodes NODES.md row + chromebox timer install
host_evidence fallback on box read registry (relay) + installed timers (browser) effective status when live probes are blind

Decisions

(D1..D7 below — all UNRESOLVED unless marked.)

IN: chromebox-watchdog ×6, cdp-relay-watchdog, agent-health + circuit breaker, ensure-node-supervision feed, host_evidence fallback. OUT: onboarding pipeline lifecycle, completion auditor, loop-health (each keeps its own owner and interviews separately).

Policy: each supervisor owns its signal — chromebox-watchdog owns CDP/egress/process failure, agent-health owns API-layer failure. Browser restart stays the shared actuator, and both keep their <2-min fresh-browser guards. No code change; this paragraph is the guard against future edits dropping either side.

Policy: stop after a failed tunnel restart. Log the egress failure loudly and page it as a tunnel fault, never as a browser fault; skip the chrome relaunch. The next 2-min run retries the tunnel. (Requires a chromebox-watchdog.sh change — implementation needs a separate explicit request.)

Keep 3 strikes / 1800s cooldown as initial values, with a review trigger: revisit if one node trips the circuit more than twice in a week (threshold too touchy) or a stuck node sits out a full cooldown under operator eyes (cooldown too long).

Registry-only: every active NODES.md row is supervised, period. The ensure script guarantees the timer follows the row, and a missing timer degrades to an honest unknown (no journal runs) rather than a wrong verdict. No second inventory.

Protocol: claim-before-edit (announce shared-file claims in chat), anchored-edits-only on shared files (never full-file rewrites, so concurrent work interleaves), re-read before staging, and prompt mine-only commits so landed work survives the next clobber.

D7. Done means — UNRESOLVED

Proposed checklist: timers on all 6 firing silent; relay/agent-health loops registry-driven with tests; ensure hook live; this doc Final.

Risks

  • Egress-down pages read as browser failures (D3).
  • Uncommitted supervisor work can be clobbered by a concurrent editor (D6).
  • c9143a5 contains unattributed peer hunks (host_evidence _covered_nodes, fleet-status test updates) — needs an amend-or-leave decision.

Validation

  • box fleet status truthful from blind shells (live-verified 6/6 ACTIVE).
  • Focused suites green (supervision, fleet, agent-health, watchdog coverage).
  • Timer firing proven via journal, not config presence.

Unresolved items

D7 unresolved. D1–D6 settled.