feat(watchers): add box stability watcher daemon, recovery guardrails, and CLI integration
- Add dedicated watchers/ project folder with box-stability-watcher.py supervisor - Monitor host load, memory, swap saturation, and crash-looping services - Implement tiered mitigations: yellow renicing, orange SIGSTOP pause with 60s grace, red shedding - Distinguish user-launched agents (allowed on desktop default socket) from automated box workloads - Wire first-class box stability CLI subcommand and top-line host status in fleet status - Harden tmux.service with cgroup memory limits to prevent OS freeze and OOM avalanches - Add 10-test unit test suite covering thresholds, safety whitelist, pause/resume, and isolation
This commit is contained in:
@@ -0,0 +1,60 @@
|
||||
# NetVM Watchers
|
||||
|
||||
Dedicated directory for background autonomous health, resource, and stability watchers on NetVM host `bl`.
|
||||
|
||||
## Components
|
||||
|
||||
- **`box-stability-watcher.py`**: Host resource supervisor and load shedder. Proactively monitors:
|
||||
- System 1m/5m/15m load averages against core counts.
|
||||
- Host RAM and Swap pressure percentages.
|
||||
- Per-process memory leaks (critical RSS thresholds for `muse-bin`, headless Chromium renderers, Python workers).
|
||||
- Rogue/leaked CPU hogs starving SSH/Tailscale.
|
||||
- Failing systemd user services trapped in tight restart loops.
|
||||
- Socket isolation violations (automated workers running on `/tmp/tmux-1000/default` instead of `/tmp/tmux-muse.sock`).
|
||||
- **`box-stability.json`**: Tunable operational thresholds, notifications, and protected process whitelist.
|
||||
- **`systemd/box-stability-watcher.service`**: Systemd user daemon running the watcher continuously with 10s evaluation ticks.
|
||||
|
||||
## Operational Tiers & Mitigations
|
||||
|
||||
| Tier | Status | Trigger Condition | Automated Action |
|
||||
| :--- | :--- | :--- | :--- |
|
||||
| **GREEN** | Normal | Load < 20, RAM < 80%, Swap < 75% | Silent monitoring. |
|
||||
| **YELLOW** | Warning | Load >= 20, RAM >= 80%, or process RSS >= 2000MB | Renice CPU hogs (+15) to preserve interactive SSH responsiveness; log warning. |
|
||||
| **ORANGE** | Critical | Load >= 35, RAM >= 90%, or process RSS >= 3000MB | **Pause (SIGSTOP)** runaway worker, record in `.state/stability-paused.json`, post alert to `646 tasks` sidechat, and allow 60s operator inspection before SIGTERM. |
|
||||
| **RED** | Emergency | Load >= 60, RAM >= 95%, or Swap >= 92% | Emergency load shedding of non-protected heavy consumers (>1500MB). |
|
||||
|
||||
## Paused Process Lifecycle (60s Grace Window)
|
||||
|
||||
When a process is paused:
|
||||
1. Sent `SIGSTOP` immediately.
|
||||
2. Recorded in `.state/stability-paused.json` with timestamp and command info.
|
||||
3. Alert posted to `646 tasks` sidechat.
|
||||
4. An operator can inspect the runtime or resume it via:
|
||||
```bash
|
||||
box stability resume <PID>
|
||||
```
|
||||
5. If unresumed after 60 seconds, the watcher automatically culls the process via `SIGTERM`.
|
||||
|
||||
## Protected Whitelist
|
||||
|
||||
The watcher will **never** terminate or renice:
|
||||
`sshd`, `tailscaled`, `tailscale`, `systemd`, `dbus-broker`, `pipewire`, `wireplumber`, `tmux` (main server), `bash`, `zsh`, `ghostty`, `alacritty`.
|
||||
|
||||
## Unified Box CLI Integration
|
||||
|
||||
```bash
|
||||
# Host stability status & active socket warnings
|
||||
box stability status
|
||||
|
||||
# Machine-readable JSON output
|
||||
box stability json
|
||||
|
||||
# Single evaluation check
|
||||
box stability check [--dry-run]
|
||||
|
||||
# Resume a paused process
|
||||
box stability resume <PID>
|
||||
|
||||
# Top-line host health indicator
|
||||
box fleet status
|
||||
```
|
||||
Reference in New Issue
Block a user