Files
box/docs/BOX-STABILITY-WATCHER.md
T
operator 0a45133d28 feat(watchers): add box stability watcher daemon, recovery guardrails, and CLI integration
- Add dedicated watchers/ project folder with box-stability-watcher.py supervisor
- Monitor host load, memory, swap saturation, and crash-looping services
- Implement tiered mitigations: yellow renicing, orange SIGSTOP pause with 60s grace, red shedding
- Distinguish user-launched agents (allowed on desktop default socket) from automated box workloads
- Wire first-class box stability CLI subcommand and top-line host status in fleet status
- Harden tmux.service with cgroup memory limits to prevent OS freeze and OOM avalanches
- Add 10-test unit test suite covering thresholds, safety whitelist, pause/resume, and isolation
2026-10-07 17:49:14 +00:00

3.3 KiB

Box Stability Watcher & Host Load Mitigation

Date: 2026-10-07
Scope: Host bl (100.123.153.75), NetVM execution stability, and preventative runaway containment.


1. Incident Post-Mortem (2026-10-07)

Symptoms

  • Host bl stopped responding over SSH and Tailscale ("went dark") at ~16:44 UTC.
  • Connections timed out during SSH banner exchange (Connection timed out during banner exchange).
  • Tailscale direct connections dropped, failing back to DERP relay nyc before dropping entirely.

Forensics & Root Cause

  1. Tmux Memory Leak & OOM Killer:
    • At 16:44:52 UTC, systemd triggered an OOM-kill on tmux.service:
      tmux.service: Consumed 5h 52min 34s CPU time ... 22.3G memory peak, 1.3G memory swap peak.
      tmux.service: Failed with result "oom-kill".
      
    • Multiple muse-bin worker processes running inside background tmux windows had accumulated 22.3 GB of memory against the host 28 GB RAM and 4 GB swap.
  2. Avalanche Load Spike (Load Avg: 515.17):
    • The unconstrained crash triggered core dump collection (systemd-coredump) and simultaneous resurrection of multiple background sessions.
    • Host 1-minute load average spiked to 515.17 (on a 16-core CPU), starving kernel network processing and dropping incoming SSH and Tailscale packets.
  3. Crash Loop Contributor:
    • Concurrently, audio-patchbay.service was stuck in a tight infinite failure loop (restarting >1,075,000 times) due to a headless GTK initialization panic, generating relentless fork/exit churn.

2. Hardening Measures Implemented

A. Dedicated Watcher Directory (Projects/NetVM/watchers/)

Created a dedicated project folder inside Projects/NetVM/ containing:

  • watchers/box-stability-watcher.py: Core stability supervisor.
  • watchers/box-stability.json: Operational configuration & thresholds.
  • watchers/README.md: Architecture and usage guide.
  • bin/box-stability-watcher.py: Symlink for operator CLI access.

B. Proactive Mitigation Tiers

  • Tier GREEN (<20 load, <80% RAM): Passive observation.
  • Tier YELLOW (20-35 load, 80-90% RAM): Renices rogue CPU hogs to nice +15 to protect interactive SSH and Tailscale responsiveness.
  • Tier ORANGE (35-60 load, >90% RAM, or process RSS >3000MB): Proactively sends SIGTERM to individual leaky worker processes (muse-bin, headless renderers) before system memory is exhausted and kernel OOM kills tmux.service.
  • Tier RED (>60 load, >95% RAM, or >92% Swap): Emergency shedder that terminates non-protected heavy consumers (>1500MB) to avert complete system lockup.

C. Systemd Hardening & Service Cleanup

  1. Disabled Runaway Service: Stopped and disabled audio-patchbay.service, instantly halting 1M+ iterations of process restart overhead.
  2. Cgroup Memory Limits on tmux.service: Configured MemoryHigh=18G and MemoryMax=22G in ~/.config/systemd/user/tmux.service so that a rogue subagent cannot consume 100% of the host RAM.
  3. Daemonized Watcher: Enabled box-stability-watcher.service as a persistent user systemd service with a 256MB memory cap and automatic restart.

3. Verification

  • Watcher unit test suite: tests/test_box_stability_watcher.py (10/10 tests passed).
  • Live evaluation: box-stability-watcher.py --status reports GREEN with host load normalized to ~2.5.