diff --git a/docs/FLEET-RESOURCE-OPTIMIZATION.md b/docs/FLEET-RESOURCE-OPTIMIZATION.md new file mode 100644 index 0000000..d676c23 --- /dev/null +++ b/docs/FLEET-RESOURCE-OPTIMIZATION.md @@ -0,0 +1,80 @@ +# Fleet Resource Optimization & Hygiene Runbook + +> **Box is the main surface.** All operator work goes through Box (box.muse-dev.online). The web UI, `box` CLI, and agents share the same API endpoints. No UI-only powers. + +**Date:** 2026-10-05 ~17:40 UTC +**Scope:** Host `bl` (100.123.153.75), NetVM headless browser fleet, user systemd timers, and resource throttling. + +--- + +## 1. Problem Summary & Root Causes + +During operational monitoring, host CPU load average was observed elevated at **4.90–5.83**. Investigation revealed three distinct contributors: + +1. **Runaway Orphaned Audio Process:** + - `PID 2661799` (`parec -d gui_sink.monitor --format=s16le --rate=48000 --channels=2`) was spinning in state `R` (runnable) at **100% of a CPU core**. + - It had accumulated **46,718 minutes (~778 hours)** of continuous CPU time since September 3 after its downstream pipe reader terminated. +2. **SwiftShader Software WebGL Overhead:** + - Headless Chromium instances running without hardware acceleration were invoking SwiftShader (`--use-angle=swiftshader-webgl`) software rasterization loops on background tab animations. +3. **Timer & Scheduler Duplication:** + - 170 individual systemd user timers (`job-auto-work-*.timer`) were active concurrently with the unified scheduler [`job-scheduler.py`](file:///home/super/Projects/NetVM/bin/job-scheduler.py), leading to excessive process spawns and duplicate dispatch checks. +4. **Lingering Scopes:** + - Previous browser watchdog relaunches occasionally killed only the Chromium child process while leaving transient systemd scopes and `netvm-cdp-relay.py` instances orphaned. + +--- + +## 2. Remediations Applied + +### A. Process & Scope Hygiene +- **Terminated Zombie Audio Streamer:** + ```bash + kill 2661798 2661799 + ``` + Immediately recovered ~1.0 off the system load average. +- **Cleared Orphaned Scopes:** + Stopped leftover transient scopes (`netvm-chrome-dev-1791158094`, `netvm-chrome-muse-1791217608`, `netvm-chrome-opm-1791217495`). +- **Hardened Watchdog Relaunch Logic:** + Updated [`bin/chromebox-watchdog.sh`](file:///home/super/Projects/NetVM/bin/chromebox-watchdog.sh) to explicitly stop existing user scopes matching `netvm-chrome-${PROFILE}-*` before starting a new transient unit. + +### B. Chromium Headless Throttling & Memory Caps +Updated [`chrome-box`](file:///home/super/Projects/chrome-box/chrome-box) for both sandboxed and native (`--no-sandbox`) headless execution paths to include: +- `--disable-software-rasterizer` & `--disable-gpu-compositing` (disables CPU-heavy 3D software rasterization). +- `--mute-audio` (prevents unnecessary audio stream synthesis). +- `--disable-dev-shm-usage` (protects `/dev/shm` exhaustion). +- `--js-flags=--max-old-space-size=768` (bounds V8 heap consumption to 768MB per profile). +- Background timer and renderer throttling enabled (`--disable-renderer-backgrounding=false`, `--disable-background-timer-throttling=false`, `--disable-backgrounding-occluded-windows=false`). + +### C. Timer Consolidation +- **Decommissioned 170 Redundant Timers:** + Archived all `job-auto-work-*.timer` and `.service` units to `~/.config/systemd/user/archive-auto-work-timers/` and disabled them. +- **Single Source of Truth:** + Retained [`job-scheduler.timer`](file:///home/super/.config/systemd/user/job-scheduler.timer) (`job-scheduler.py run` on `*:0/5`), which evaluates cron schedules in [`jobs/*.json`](file:///home/super/Projects/NetVM/jobs/) on a unified schedule. User timers dropped from **190 to 20**. + +--- + +## 3. Verification & Fleet Health + +Verification was conducted across all operational layers: + +1. **Fleet Status:** + ```bash + super fleet status + ``` + All 6 nodes (`muse`, `pip`, `646`, `opm`, `def`, `dev`) are **`● ACTIVE`** with sub-10ms CDP roundtrip latencies. +2. **Periodic Health Checks:** + ```bash + tail -n 15 /tmp/agent-health.log + ``` + Consecutive runs of `agent-health.service` confirmed 100% OK status across all nodes. +3. **Scheduler Status:** + ```bash + python3 bin/job-scheduler.py status + ``` + Unified scheduler runs cleanly every 5 minutes and updates `job-scheduler-state.json`. +4. **Log Retention:** + ```bash + bin/retention-run-rotations.sh + ``` + Completed with return code `0`; archives passed gzip integrity and checksum validations. +5. **System Load:** + Load average dropped from **~5.8** to **~2.8–3.2** (~18% utilization across 16 logical cores).