57 lines
3.3 KiB
Markdown
57 lines
3.3 KiB
Markdown
|
|
# Box Stability Watcher & Host Load Mitigation
|
||
|
|
|
||
|
|
**Date:** 2026-10-07
|
||
|
|
**Scope:** Host `bl` (100.123.153.75), NetVM execution stability, and preventative runaway containment.
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## 1. Incident Post-Mortem (2026-10-07)
|
||
|
|
|
||
|
|
### Symptoms
|
||
|
|
- Host `bl` stopped responding over SSH and Tailscale ("went dark") at ~16:44 UTC.
|
||
|
|
- Connections timed out during SSH banner exchange (`Connection timed out during banner exchange`).
|
||
|
|
- Tailscale direct connections dropped, failing back to DERP relay `nyc` before dropping entirely.
|
||
|
|
|
||
|
|
### Forensics & Root Cause
|
||
|
|
1. **Tmux Memory Leak & OOM Killer:**
|
||
|
|
- At 16:44:52 UTC, `systemd` triggered an OOM-kill on `tmux.service`:
|
||
|
|
```
|
||
|
|
tmux.service: Consumed 5h 52min 34s CPU time ... 22.3G memory peak, 1.3G memory swap peak.
|
||
|
|
tmux.service: Failed with result "oom-kill".
|
||
|
|
```
|
||
|
|
- Multiple `muse-bin` worker processes running inside background tmux windows had accumulated 22.3 GB of memory against the host 28 GB RAM and 4 GB swap.
|
||
|
|
2. **Avalanche Load Spike (Load Avg: 515.17):**
|
||
|
|
- The unconstrained crash triggered core dump collection (`systemd-coredump`) and simultaneous resurrection of multiple background sessions.
|
||
|
|
- Host 1-minute load average spiked to **515.17** (on a 16-core CPU), starving kernel network processing and dropping incoming SSH and Tailscale packets.
|
||
|
|
3. **Crash Loop Contributor:**
|
||
|
|
- Concurrently, `audio-patchbay.service` was stuck in a tight infinite failure loop (restarting >1,075,000 times) due to a headless GTK initialization panic, generating relentless fork/exit churn.
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## 2. Hardening Measures Implemented
|
||
|
|
|
||
|
|
### A. Dedicated Watcher Directory (`Projects/NetVM/watchers/`)
|
||
|
|
Created a dedicated project folder inside `Projects/NetVM/` containing:
|
||
|
|
- `watchers/box-stability-watcher.py`: Core stability supervisor.
|
||
|
|
- `watchers/box-stability.json`: Operational configuration & thresholds.
|
||
|
|
- `watchers/README.md`: Architecture and usage guide.
|
||
|
|
- `bin/box-stability-watcher.py`: Symlink for operator CLI access.
|
||
|
|
|
||
|
|
### B. Proactive Mitigation Tiers
|
||
|
|
- **Tier GREEN (<20 load, <80% RAM):** Passive observation.
|
||
|
|
- **Tier YELLOW (20-35 load, 80-90% RAM):** Renices rogue CPU hogs to `nice +15` to protect interactive SSH and Tailscale responsiveness.
|
||
|
|
- **Tier ORANGE (35-60 load, >90% RAM, or process RSS >3000MB):** Proactively sends `SIGTERM` to individual leaky worker processes (`muse-bin`, headless renderers) before system memory is exhausted and kernel OOM kills `tmux.service`.
|
||
|
|
- **Tier RED (>60 load, >95% RAM, or >92% Swap):** Emergency shedder that terminates non-protected heavy consumers (>1500MB) to avert complete system lockup.
|
||
|
|
|
||
|
|
### C. Systemd Hardening & Service Cleanup
|
||
|
|
1. **Disabled Runaway Service:** Stopped and disabled `audio-patchbay.service`, instantly halting 1M+ iterations of process restart overhead.
|
||
|
|
2. **Cgroup Memory Limits on `tmux.service`:** Configured `MemoryHigh=18G` and `MemoryMax=22G` in `~/.config/systemd/user/tmux.service` so that a rogue subagent cannot consume 100% of the host RAM.
|
||
|
|
3. **Daemonized Watcher:** Enabled `box-stability-watcher.service` as a persistent user systemd service with a 256MB memory cap and automatic restart.
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## 3. Verification
|
||
|
|
|
||
|
|
- Watcher unit test suite: `tests/test_box_stability_watcher.py` (10/10 tests passed).
|
||
|
|
- Live evaluation: `box-stability-watcher.py --status` reports `GREEN` with host load normalized to ~2.5.
|