feat(watchers): add box stability watcher daemon, recovery guardrails, and CLI integration
- Add dedicated watchers/ project folder with box-stability-watcher.py supervisor - Monitor host load, memory, swap saturation, and crash-looping services - Implement tiered mitigations: yellow renicing, orange SIGSTOP pause with 60s grace, red shedding - Distinguish user-launched agents (allowed on desktop default socket) from automated box workloads - Wire first-class box stability CLI subcommand and top-line host status in fleet status - Harden tmux.service with cgroup memory limits to prevent OS freeze and OOM avalanches - Add 10-test unit test suite covering thresholds, safety whitelist, pause/resume, and isolation
This commit is contained in:
@@ -0,0 +1,56 @@
|
||||
# Box Stability Watcher & Host Load Mitigation
|
||||
|
||||
**Date:** 2026-10-07
|
||||
**Scope:** Host `bl` (100.123.153.75), NetVM execution stability, and preventative runaway containment.
|
||||
|
||||
---
|
||||
|
||||
## 1. Incident Post-Mortem (2026-10-07)
|
||||
|
||||
### Symptoms
|
||||
- Host `bl` stopped responding over SSH and Tailscale ("went dark") at ~16:44 UTC.
|
||||
- Connections timed out during SSH banner exchange (`Connection timed out during banner exchange`).
|
||||
- Tailscale direct connections dropped, failing back to DERP relay `nyc` before dropping entirely.
|
||||
|
||||
### Forensics & Root Cause
|
||||
1. **Tmux Memory Leak & OOM Killer:**
|
||||
- At 16:44:52 UTC, `systemd` triggered an OOM-kill on `tmux.service`:
|
||||
```
|
||||
tmux.service: Consumed 5h 52min 34s CPU time ... 22.3G memory peak, 1.3G memory swap peak.
|
||||
tmux.service: Failed with result "oom-kill".
|
||||
```
|
||||
- Multiple `muse-bin` worker processes running inside background tmux windows had accumulated 22.3 GB of memory against the host 28 GB RAM and 4 GB swap.
|
||||
2. **Avalanche Load Spike (Load Avg: 515.17):**
|
||||
- The unconstrained crash triggered core dump collection (`systemd-coredump`) and simultaneous resurrection of multiple background sessions.
|
||||
- Host 1-minute load average spiked to **515.17** (on a 16-core CPU), starving kernel network processing and dropping incoming SSH and Tailscale packets.
|
||||
3. **Crash Loop Contributor:**
|
||||
- Concurrently, `audio-patchbay.service` was stuck in a tight infinite failure loop (restarting >1,075,000 times) due to a headless GTK initialization panic, generating relentless fork/exit churn.
|
||||
|
||||
---
|
||||
|
||||
## 2. Hardening Measures Implemented
|
||||
|
||||
### A. Dedicated Watcher Directory (`Projects/NetVM/watchers/`)
|
||||
Created a dedicated project folder inside `Projects/NetVM/` containing:
|
||||
- `watchers/box-stability-watcher.py`: Core stability supervisor.
|
||||
- `watchers/box-stability.json`: Operational configuration & thresholds.
|
||||
- `watchers/README.md`: Architecture and usage guide.
|
||||
- `bin/box-stability-watcher.py`: Symlink for operator CLI access.
|
||||
|
||||
### B. Proactive Mitigation Tiers
|
||||
- **Tier GREEN (<20 load, <80% RAM):** Passive observation.
|
||||
- **Tier YELLOW (20-35 load, 80-90% RAM):** Renices rogue CPU hogs to `nice +15` to protect interactive SSH and Tailscale responsiveness.
|
||||
- **Tier ORANGE (35-60 load, >90% RAM, or process RSS >3000MB):** Proactively sends `SIGTERM` to individual leaky worker processes (`muse-bin`, headless renderers) before system memory is exhausted and kernel OOM kills `tmux.service`.
|
||||
- **Tier RED (>60 load, >95% RAM, or >92% Swap):** Emergency shedder that terminates non-protected heavy consumers (>1500MB) to avert complete system lockup.
|
||||
|
||||
### C. Systemd Hardening & Service Cleanup
|
||||
1. **Disabled Runaway Service:** Stopped and disabled `audio-patchbay.service`, instantly halting 1M+ iterations of process restart overhead.
|
||||
2. **Cgroup Memory Limits on `tmux.service`:** Configured `MemoryHigh=18G` and `MemoryMax=22G` in `~/.config/systemd/user/tmux.service` so that a rogue subagent cannot consume 100% of the host RAM.
|
||||
3. **Daemonized Watcher:** Enabled `box-stability-watcher.service` as a persistent user systemd service with a 256MB memory cap and automatic restart.
|
||||
|
||||
---
|
||||
|
||||
## 3. Verification
|
||||
|
||||
- Watcher unit test suite: `tests/test_box_stability_watcher.py` (10/10 tests passed).
|
||||
- Live evaluation: `box-stability-watcher.py --status` reports `GREEN` with host load normalized to ~2.5.
|
||||
Reference in New Issue
Block a user