Files
box/docs/BOX-STABILITY-WATCHER.md
T
operator 0a45133d28 feat(watchers): add box stability watcher daemon, recovery guardrails, and CLI integration
- Add dedicated watchers/ project folder with box-stability-watcher.py supervisor
- Monitor host load, memory, swap saturation, and crash-looping services
- Implement tiered mitigations: yellow renicing, orange SIGSTOP pause with 60s grace, red shedding
- Distinguish user-launched agents (allowed on desktop default socket) from automated box workloads
- Wire first-class box stability CLI subcommand and top-line host status in fleet status
- Harden tmux.service with cgroup memory limits to prevent OS freeze and OOM avalanches
- Add 10-test unit test suite covering thresholds, safety whitelist, pause/resume, and isolation
2026-10-07 17:49:14 +00:00

57 lines
3.3 KiB
Markdown

# Box Stability Watcher & Host Load Mitigation
**Date:** 2026-10-07
**Scope:** Host `bl` (100.123.153.75), NetVM execution stability, and preventative runaway containment.
---
## 1. Incident Post-Mortem (2026-10-07)
### Symptoms
- Host `bl` stopped responding over SSH and Tailscale ("went dark") at ~16:44 UTC.
- Connections timed out during SSH banner exchange (`Connection timed out during banner exchange`).
- Tailscale direct connections dropped, failing back to DERP relay `nyc` before dropping entirely.
### Forensics & Root Cause
1. **Tmux Memory Leak & OOM Killer:**
- At 16:44:52 UTC, `systemd` triggered an OOM-kill on `tmux.service`:
```
tmux.service: Consumed 5h 52min 34s CPU time ... 22.3G memory peak, 1.3G memory swap peak.
tmux.service: Failed with result "oom-kill".
```
- Multiple `muse-bin` worker processes running inside background tmux windows had accumulated 22.3 GB of memory against the host 28 GB RAM and 4 GB swap.
2. **Avalanche Load Spike (Load Avg: 515.17):**
- The unconstrained crash triggered core dump collection (`systemd-coredump`) and simultaneous resurrection of multiple background sessions.
- Host 1-minute load average spiked to **515.17** (on a 16-core CPU), starving kernel network processing and dropping incoming SSH and Tailscale packets.
3. **Crash Loop Contributor:**
- Concurrently, `audio-patchbay.service` was stuck in a tight infinite failure loop (restarting >1,075,000 times) due to a headless GTK initialization panic, generating relentless fork/exit churn.
---
## 2. Hardening Measures Implemented
### A. Dedicated Watcher Directory (`Projects/NetVM/watchers/`)
Created a dedicated project folder inside `Projects/NetVM/` containing:
- `watchers/box-stability-watcher.py`: Core stability supervisor.
- `watchers/box-stability.json`: Operational configuration & thresholds.
- `watchers/README.md`: Architecture and usage guide.
- `bin/box-stability-watcher.py`: Symlink for operator CLI access.
### B. Proactive Mitigation Tiers
- **Tier GREEN (<20 load, <80% RAM):** Passive observation.
- **Tier YELLOW (20-35 load, 80-90% RAM):** Renices rogue CPU hogs to `nice +15` to protect interactive SSH and Tailscale responsiveness.
- **Tier ORANGE (35-60 load, >90% RAM, or process RSS >3000MB):** Proactively sends `SIGTERM` to individual leaky worker processes (`muse-bin`, headless renderers) before system memory is exhausted and kernel OOM kills `tmux.service`.
- **Tier RED (>60 load, >95% RAM, or >92% Swap):** Emergency shedder that terminates non-protected heavy consumers (>1500MB) to avert complete system lockup.
### C. Systemd Hardening & Service Cleanup
1. **Disabled Runaway Service:** Stopped and disabled `audio-patchbay.service`, instantly halting 1M+ iterations of process restart overhead.
2. **Cgroup Memory Limits on `tmux.service`:** Configured `MemoryHigh=18G` and `MemoryMax=22G` in `~/.config/systemd/user/tmux.service` so that a rogue subagent cannot consume 100% of the host RAM.
3. **Daemonized Watcher:** Enabled `box-stability-watcher.service` as a persistent user systemd service with a 256MB memory cap and automatic restart.
---
## 3. Verification
- Watcher unit test suite: `tests/test_box_stability_watcher.py` (10/10 tests passed).
- Live evaluation: `box-stability-watcher.py --status` reports `GREEN` with host load normalized to ~2.5.