Files
box/watchers/README.md
T
operator 0a45133d28 feat(watchers): add box stability watcher daemon, recovery guardrails, and CLI integration
- Add dedicated watchers/ project folder with box-stability-watcher.py supervisor
- Monitor host load, memory, swap saturation, and crash-looping services
- Implement tiered mitigations: yellow renicing, orange SIGSTOP pause with 60s grace, red shedding
- Distinguish user-launched agents (allowed on desktop default socket) from automated box workloads
- Wire first-class box stability CLI subcommand and top-line host status in fleet status
- Harden tmux.service with cgroup memory limits to prevent OS freeze and OOM avalanches
- Add 10-test unit test suite covering thresholds, safety whitelist, pause/resume, and isolation
2026-10-07 17:49:14 +00:00

2.6 KiB

NetVM Watchers

Dedicated directory for background autonomous health, resource, and stability watchers on NetVM host bl.

Components

  • box-stability-watcher.py: Host resource supervisor and load shedder. Proactively monitors:
    • System 1m/5m/15m load averages against core counts.
    • Host RAM and Swap pressure percentages.
    • Per-process memory leaks (critical RSS thresholds for muse-bin, headless Chromium renderers, Python workers).
    • Rogue/leaked CPU hogs starving SSH/Tailscale.
    • Failing systemd user services trapped in tight restart loops.
    • Socket isolation violations (automated workers running on /tmp/tmux-1000/default instead of /tmp/tmux-muse.sock).
  • box-stability.json: Tunable operational thresholds, notifications, and protected process whitelist.
  • systemd/box-stability-watcher.service: Systemd user daemon running the watcher continuously with 10s evaluation ticks.

Operational Tiers & Mitigations

Tier Status Trigger Condition Automated Action
GREEN Normal Load < 20, RAM < 80%, Swap < 75% Silent monitoring.
YELLOW Warning Load >= 20, RAM >= 80%, or process RSS >= 2000MB Renice CPU hogs (+15) to preserve interactive SSH responsiveness; log warning.
ORANGE Critical Load >= 35, RAM >= 90%, or process RSS >= 3000MB Pause (SIGSTOP) runaway worker, record in .state/stability-paused.json, post alert to 646 tasks sidechat, and allow 60s operator inspection before SIGTERM.
RED Emergency Load >= 60, RAM >= 95%, or Swap >= 92% Emergency load shedding of non-protected heavy consumers (>1500MB).

Paused Process Lifecycle (60s Grace Window)

When a process is paused:

  1. Sent SIGSTOP immediately.
  2. Recorded in .state/stability-paused.json with timestamp and command info.
  3. Alert posted to 646 tasks sidechat.
  4. An operator can inspect the runtime or resume it via:
    box stability resume <PID>
    
  5. If unresumed after 60 seconds, the watcher automatically culls the process via SIGTERM.

Protected Whitelist

The watcher will never terminate or renice: sshd, tailscaled, tailscale, systemd, dbus-broker, pipewire, wireplumber, tmux (main server), bash, zsh, ghostty, alacritty.

Unified Box CLI Integration

# Host stability status & active socket warnings
box stability status

# Machine-readable JSON output
box stability json

# Single evaluation check
box stability check [--dry-run]

# Resume a paused process
box stability resume <PID>

# Top-line host health indicator
box fleet status