0a45133d28
- Add dedicated watchers/ project folder with box-stability-watcher.py supervisor - Monitor host load, memory, swap saturation, and crash-looping services - Implement tiered mitigations: yellow renicing, orange SIGSTOP pause with 60s grace, red shedding - Distinguish user-launched agents (allowed on desktop default socket) from automated box workloads - Wire first-class box stability CLI subcommand and top-line host status in fleet status - Harden tmux.service with cgroup memory limits to prevent OS freeze and OOM avalanches - Add 10-test unit test suite covering thresholds, safety whitelist, pause/resume, and isolation
3.3 KiB
3.3 KiB
Box Stability Watcher & Host Load Mitigation
Date: 2026-10-07
Scope: Host bl (100.123.153.75), NetVM execution stability, and preventative runaway containment.
1. Incident Post-Mortem (2026-10-07)
Symptoms
- Host
blstopped responding over SSH and Tailscale ("went dark") at ~16:44 UTC. - Connections timed out during SSH banner exchange (
Connection timed out during banner exchange). - Tailscale direct connections dropped, failing back to DERP relay
nycbefore dropping entirely.
Forensics & Root Cause
- Tmux Memory Leak & OOM Killer:
- At 16:44:52 UTC,
systemdtriggered an OOM-kill ontmux.service:tmux.service: Consumed 5h 52min 34s CPU time ... 22.3G memory peak, 1.3G memory swap peak. tmux.service: Failed with result "oom-kill". - Multiple
muse-binworker processes running inside background tmux windows had accumulated 22.3 GB of memory against the host 28 GB RAM and 4 GB swap.
- At 16:44:52 UTC,
- Avalanche Load Spike (Load Avg: 515.17):
- The unconstrained crash triggered core dump collection (
systemd-coredump) and simultaneous resurrection of multiple background sessions. - Host 1-minute load average spiked to 515.17 (on a 16-core CPU), starving kernel network processing and dropping incoming SSH and Tailscale packets.
- The unconstrained crash triggered core dump collection (
- Crash Loop Contributor:
- Concurrently,
audio-patchbay.servicewas stuck in a tight infinite failure loop (restarting >1,075,000 times) due to a headless GTK initialization panic, generating relentless fork/exit churn.
- Concurrently,
2. Hardening Measures Implemented
A. Dedicated Watcher Directory (Projects/NetVM/watchers/)
Created a dedicated project folder inside Projects/NetVM/ containing:
watchers/box-stability-watcher.py: Core stability supervisor.watchers/box-stability.json: Operational configuration & thresholds.watchers/README.md: Architecture and usage guide.bin/box-stability-watcher.py: Symlink for operator CLI access.
B. Proactive Mitigation Tiers
- Tier GREEN (<20 load, <80% RAM): Passive observation.
- Tier YELLOW (20-35 load, 80-90% RAM): Renices rogue CPU hogs to
nice +15to protect interactive SSH and Tailscale responsiveness. - Tier ORANGE (35-60 load, >90% RAM, or process RSS >3000MB): Proactively sends
SIGTERMto individual leaky worker processes (muse-bin, headless renderers) before system memory is exhausted and kernel OOM killstmux.service. - Tier RED (>60 load, >95% RAM, or >92% Swap): Emergency shedder that terminates non-protected heavy consumers (>1500MB) to avert complete system lockup.
C. Systemd Hardening & Service Cleanup
- Disabled Runaway Service: Stopped and disabled
audio-patchbay.service, instantly halting 1M+ iterations of process restart overhead. - Cgroup Memory Limits on
tmux.service: ConfiguredMemoryHigh=18GandMemoryMax=22Gin~/.config/systemd/user/tmux.serviceso that a rogue subagent cannot consume 100% of the host RAM. - Daemonized Watcher: Enabled
box-stability-watcher.serviceas a persistent user systemd service with a 256MB memory cap and automatic restart.
3. Verification
- Watcher unit test suite:
tests/test_box_stability_watcher.py(10/10 tests passed). - Live evaluation:
box-stability-watcher.py --statusreportsGREENwith host load normalized to ~2.5.