0a45133d28
- Add dedicated watchers/ project folder with box-stability-watcher.py supervisor - Monitor host load, memory, swap saturation, and crash-looping services - Implement tiered mitigations: yellow renicing, orange SIGSTOP pause with 60s grace, red shedding - Distinguish user-launched agents (allowed on desktop default socket) from automated box workloads - Wire first-class box stability CLI subcommand and top-line host status in fleet status - Harden tmux.service with cgroup memory limits to prevent OS freeze and OOM avalanches - Add 10-test unit test suite covering thresholds, safety whitelist, pause/resume, and isolation
61 lines
2.6 KiB
Markdown
61 lines
2.6 KiB
Markdown
# NetVM Watchers
|
|
|
|
Dedicated directory for background autonomous health, resource, and stability watchers on NetVM host `bl`.
|
|
|
|
## Components
|
|
|
|
- **`box-stability-watcher.py`**: Host resource supervisor and load shedder. Proactively monitors:
|
|
- System 1m/5m/15m load averages against core counts.
|
|
- Host RAM and Swap pressure percentages.
|
|
- Per-process memory leaks (critical RSS thresholds for `muse-bin`, headless Chromium renderers, Python workers).
|
|
- Rogue/leaked CPU hogs starving SSH/Tailscale.
|
|
- Failing systemd user services trapped in tight restart loops.
|
|
- Socket isolation violations (automated workers running on `/tmp/tmux-1000/default` instead of `/tmp/tmux-muse.sock`).
|
|
- **`box-stability.json`**: Tunable operational thresholds, notifications, and protected process whitelist.
|
|
- **`systemd/box-stability-watcher.service`**: Systemd user daemon running the watcher continuously with 10s evaluation ticks.
|
|
|
|
## Operational Tiers & Mitigations
|
|
|
|
| Tier | Status | Trigger Condition | Automated Action |
|
|
| :--- | :--- | :--- | :--- |
|
|
| **GREEN** | Normal | Load < 20, RAM < 80%, Swap < 75% | Silent monitoring. |
|
|
| **YELLOW** | Warning | Load >= 20, RAM >= 80%, or process RSS >= 2000MB | Renice CPU hogs (+15) to preserve interactive SSH responsiveness; log warning. |
|
|
| **ORANGE** | Critical | Load >= 35, RAM >= 90%, or process RSS >= 3000MB | **Pause (SIGSTOP)** runaway worker, record in `.state/stability-paused.json`, post alert to `646 tasks` sidechat, and allow 60s operator inspection before SIGTERM. |
|
|
| **RED** | Emergency | Load >= 60, RAM >= 95%, or Swap >= 92% | Emergency load shedding of non-protected heavy consumers (>1500MB). |
|
|
|
|
## Paused Process Lifecycle (60s Grace Window)
|
|
|
|
When a process is paused:
|
|
1. Sent `SIGSTOP` immediately.
|
|
2. Recorded in `.state/stability-paused.json` with timestamp and command info.
|
|
3. Alert posted to `646 tasks` sidechat.
|
|
4. An operator can inspect the runtime or resume it via:
|
|
```bash
|
|
box stability resume <PID>
|
|
```
|
|
5. If unresumed after 60 seconds, the watcher automatically culls the process via `SIGTERM`.
|
|
|
|
## Protected Whitelist
|
|
|
|
The watcher will **never** terminate or renice:
|
|
`sshd`, `tailscaled`, `tailscale`, `systemd`, `dbus-broker`, `pipewire`, `wireplumber`, `tmux` (main server), `bash`, `zsh`, `ghostty`, `alacritty`.
|
|
|
|
## Unified Box CLI Integration
|
|
|
|
```bash
|
|
# Host stability status & active socket warnings
|
|
box stability status
|
|
|
|
# Machine-readable JSON output
|
|
box stability json
|
|
|
|
# Single evaluation check
|
|
box stability check [--dry-run]
|
|
|
|
# Resume a paused process
|
|
box stability resume <PID>
|
|
|
|
# Top-line host health indicator
|
|
box fleet status
|
|
```
|