feat(watchers): add box stability watcher daemon, recovery guardrails, and CLI integration

- Add dedicated watchers/ project folder with box-stability-watcher.py supervisor
- Monitor host load, memory, swap saturation, and crash-looping services
- Implement tiered mitigations: yellow renicing, orange SIGSTOP pause with 60s grace, red shedding
- Distinguish user-launched agents (allowed on desktop default socket) from automated box workloads
- Wire first-class box stability CLI subcommand and top-line host status in fleet status
- Harden tmux.service with cgroup memory limits to prevent OS freeze and OOM avalanches
- Add 10-test unit test suite covering thresholds, safety whitelist, pause/resume, and isolation
This commit is contained in:
operator
2026-10-07 17:49:14 +00:00
parent c11d1d83ae
commit 0a45133d28
8 changed files with 1616 additions and 22 deletions
+60
View File
@@ -0,0 +1,60 @@
# NetVM Watchers
Dedicated directory for background autonomous health, resource, and stability watchers on NetVM host `bl`.
## Components
- **`box-stability-watcher.py`**: Host resource supervisor and load shedder. Proactively monitors:
- System 1m/5m/15m load averages against core counts.
- Host RAM and Swap pressure percentages.
- Per-process memory leaks (critical RSS thresholds for `muse-bin`, headless Chromium renderers, Python workers).
- Rogue/leaked CPU hogs starving SSH/Tailscale.
- Failing systemd user services trapped in tight restart loops.
- Socket isolation violations (automated workers running on `/tmp/tmux-1000/default` instead of `/tmp/tmux-muse.sock`).
- **`box-stability.json`**: Tunable operational thresholds, notifications, and protected process whitelist.
- **`systemd/box-stability-watcher.service`**: Systemd user daemon running the watcher continuously with 10s evaluation ticks.
## Operational Tiers & Mitigations
| Tier | Status | Trigger Condition | Automated Action |
| :--- | :--- | :--- | :--- |
| **GREEN** | Normal | Load < 20, RAM < 80%, Swap < 75% | Silent monitoring. |
| **YELLOW** | Warning | Load >= 20, RAM >= 80%, or process RSS >= 2000MB | Renice CPU hogs (+15) to preserve interactive SSH responsiveness; log warning. |
| **ORANGE** | Critical | Load >= 35, RAM >= 90%, or process RSS >= 3000MB | **Pause (SIGSTOP)** runaway worker, record in `.state/stability-paused.json`, post alert to `646 tasks` sidechat, and allow 60s operator inspection before SIGTERM. |
| **RED** | Emergency | Load >= 60, RAM >= 95%, or Swap >= 92% | Emergency load shedding of non-protected heavy consumers (>1500MB). |
## Paused Process Lifecycle (60s Grace Window)
When a process is paused:
1. Sent `SIGSTOP` immediately.
2. Recorded in `.state/stability-paused.json` with timestamp and command info.
3. Alert posted to `646 tasks` sidechat.
4. An operator can inspect the runtime or resume it via:
```bash
box stability resume <PID>
```
5. If unresumed after 60 seconds, the watcher automatically culls the process via `SIGTERM`.
## Protected Whitelist
The watcher will **never** terminate or renice:
`sshd`, `tailscaled`, `tailscale`, `systemd`, `dbus-broker`, `pipewire`, `wireplumber`, `tmux` (main server), `bash`, `zsh`, `ghostty`, `alacritty`.
## Unified Box CLI Integration
```bash
# Host stability status & active socket warnings
box stability status
# Machine-readable JSON output
box stability json
# Single evaluation check
box stability check [--dry-run]
# Resume a paused process
box stability resume <PID>
# Top-line host health indicator
box fleet status
```