# Box Stability Watcher & Host Load Mitigation **Date:** 2026-10-07 **Scope:** Host `bl` (100.123.153.75), NetVM execution stability, and preventative runaway containment. --- ## 1. Incident Post-Mortem (2026-10-07) ### Symptoms - Host `bl` stopped responding over SSH and Tailscale ("went dark") at ~16:44 UTC. - Connections timed out during SSH banner exchange (`Connection timed out during banner exchange`). - Tailscale direct connections dropped, failing back to DERP relay `nyc` before dropping entirely. ### Forensics & Root Cause 1. **Tmux Memory Leak & OOM Killer:** - At 16:44:52 UTC, `systemd` triggered an OOM-kill on `tmux.service`: ``` tmux.service: Consumed 5h 52min 34s CPU time ... 22.3G memory peak, 1.3G memory swap peak. tmux.service: Failed with result "oom-kill". ``` - Multiple `muse-bin` worker processes running inside background tmux windows had accumulated 22.3 GB of memory against the host 28 GB RAM and 4 GB swap. 2. **Avalanche Load Spike (Load Avg: 515.17):** - The unconstrained crash triggered core dump collection (`systemd-coredump`) and simultaneous resurrection of multiple background sessions. - Host 1-minute load average spiked to **515.17** (on a 16-core CPU), starving kernel network processing and dropping incoming SSH and Tailscale packets. 3. **Crash Loop Contributor:** - Concurrently, `audio-patchbay.service` was stuck in a tight infinite failure loop (restarting >1,075,000 times) due to a headless GTK initialization panic, generating relentless fork/exit churn. --- ## 2. Hardening Measures Implemented ### A. Dedicated Watcher Directory (`Projects/NetVM/watchers/`) Created a dedicated project folder inside `Projects/NetVM/` containing: - `watchers/box-stability-watcher.py`: Core stability supervisor. - `watchers/box-stability.json`: Operational configuration & thresholds. - `watchers/README.md`: Architecture and usage guide. - `bin/box-stability-watcher.py`: Symlink for operator CLI access. ### B. Proactive Mitigation Tiers - **Tier GREEN (<20 load, <80% RAM):** Passive observation. - **Tier YELLOW (20-35 load, 80-90% RAM):** Renices rogue CPU hogs to `nice +15` to protect interactive SSH and Tailscale responsiveness. - **Tier ORANGE (35-60 load, >90% RAM, or process RSS >3000MB):** Proactively sends `SIGTERM` to individual leaky worker processes (`muse-bin`, headless renderers) before system memory is exhausted and kernel OOM kills `tmux.service`. - **Tier RED (>60 load, >95% RAM, or >92% Swap):** Emergency shedder that terminates non-protected heavy consumers (>1500MB) to avert complete system lockup. ### C. Systemd Hardening & Service Cleanup 1. **Disabled Runaway Service:** Stopped and disabled `audio-patchbay.service`, instantly halting 1M+ iterations of process restart overhead. 2. **Cgroup Memory Limits on `tmux.service`:** Configured `MemoryHigh=18G` and `MemoryMax=22G` in `~/.config/systemd/user/tmux.service` so that a rogue subagent cannot consume 100% of the host RAM. 3. **Daemonized Watcher:** Enabled `box-stability-watcher.service` as a persistent user systemd service with a 256MB memory cap and automatic restart. --- ## 3. Verification - Watcher unit test suite: `tests/test_box_stability_watcher.py` (10/10 tests passed). - Live evaluation: `box-stability-watcher.py --status` reports `GREEN` with host load normalized to ~2.5.