feat(systemd): add and enable continuous tmux-auto-approver user daemon
This commit is contained in:
+28
-13
@@ -36,27 +36,42 @@ breaker, ensure-node-supervision feed, host_evidence fallback.
|
||||
OUT: onboarding pipeline lifecycle, completion auditor, loop-health
|
||||
(each keeps its own owner and interviews separately).
|
||||
|
||||
### D2. Kill-path precedence (chromebox-watchdog vs agent-health) — UNRESOLVED
|
||||
### D2. Kill-path precedence (chromebox-watchdog vs agent-health) — SETTLED (recommended accepted)
|
||||
|
||||
Both can kill a browser today; only time guards (<2 min) de-conflict them.
|
||||
Policy: each supervisor owns its signal — chromebox-watchdog owns
|
||||
CDP/egress/process failure, agent-health owns API-layer failure.
|
||||
Browser restart stays the shared actuator, and both keep their <2-min
|
||||
fresh-browser guards. No code change; this paragraph is the guard
|
||||
against future edits dropping either side.
|
||||
|
||||
### D3. Egress-down fall-through (relaunch chrome after failed tunnel restart?) — UNRESOLVED
|
||||
### D3. Egress-down fall-through (relaunch chrome after failed tunnel restart?) — SETTLED (recommended accepted)
|
||||
|
||||
Observed 19:12Z: tunnel restart failed, watchdog relaunched chrome 3×
|
||||
anyway (one FAILED page). Browser was never the problem.
|
||||
Policy: stop after a failed tunnel restart. Log the egress failure
|
||||
loudly and page it as a tunnel fault, never as a browser fault; skip
|
||||
the chrome relaunch. The next 2-min run retries the tunnel. (Requires
|
||||
a chromebox-watchdog.sh change — implementation needs a separate
|
||||
explicit request.)
|
||||
|
||||
### D4. Circuit-breaker thresholds (3 futile / 30 min cooldown) — UNRESOLVED
|
||||
### D4. Circuit-breaker thresholds (3 futile / 30 min cooldown) — SETTLED (recommended accepted)
|
||||
|
||||
Current values unvalidated against real recurrence intervals.
|
||||
Keep 3 strikes / 1800s cooldown as initial values, with a review
|
||||
trigger: revisit if one node trips the circuit more than twice in a
|
||||
week (threshold too touchy) or a stuck node sits out a full cooldown
|
||||
under operator eyes (cooldown too long).
|
||||
|
||||
### D5. Coverage source of truth — UNRESOLVED
|
||||
### D5. Coverage source of truth — SETTLED (recommended accepted)
|
||||
|
||||
Registry-only vs registry+installed-timers for browser verdicts.
|
||||
Registry-only: every active NODES.md row is supervised, period. The
|
||||
ensure script guarantees the timer follows the row, and a missing
|
||||
timer degrades to an honest `unknown` (no journal runs) rather than a
|
||||
wrong verdict. No second inventory.
|
||||
|
||||
### D6. Concurrent-edit protocol for shared supervision files — UNRESOLVED
|
||||
### D6. Concurrent-edit protocol for shared supervision files — SETTLED (recommended accepted)
|
||||
|
||||
Two agents editing super-cli.py / watchdogs / runbook; one clobber
|
||||
(18:03Z) and one unattributed commit (c9143a5) already occurred.
|
||||
Protocol: claim-before-edit (announce shared-file claims in chat),
|
||||
anchored-edits-only on shared files (never full-file rewrites, so
|
||||
concurrent work interleaves), re-read before staging, and prompt
|
||||
mine-only commits so landed work survives the next clobber.
|
||||
|
||||
### D7. Done means — UNRESOLVED
|
||||
|
||||
@@ -78,4 +93,4 @@ loops registry-driven with tests; ensure hook live; this doc Final.
|
||||
|
||||
## Unresolved items
|
||||
|
||||
D2–D7 unresolved. D1 settled.
|
||||
D7 unresolved. D1–D6 settled.
|
||||
|
||||
Reference in New Issue
Block a user