Files
box/docs/FLEET-RESOURCE-OPTIMIZATION.md

81 lines
4.3 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Fleet Resource Optimization & Hygiene Runbook
> **Box is the main surface.** All operator work goes through Box (box.muse-dev.online). The web UI, `box` CLI, and agents share the same API endpoints. No UI-only powers.
**Date:** 2026-10-05 ~17:40 UTC
**Scope:** Host `bl` (100.123.153.75), NetVM headless browser fleet, user systemd timers, and resource throttling.
---
## 1. Problem Summary & Root Causes
During operational monitoring, host CPU load average was observed elevated at **4.90–5.83**. Investigation revealed three distinct contributors:
1. **Runaway Orphaned Audio Process:**
- `PID 2661799` (`parec -d gui_sink.monitor --format=s16le --rate=48000 --channels=2`) was spinning in state `R` (runnable) at **100% of a CPU core**.
- It had accumulated **46,718 minutes (~778 hours)** of continuous CPU time since September 3 after its downstream pipe reader terminated.
2. **SwiftShader Software WebGL Overhead:**
- Headless Chromium instances running without hardware acceleration were invoking SwiftShader (`--use-angle=swiftshader-webgl`) software rasterization loops on background tab animations.
3. **Timer & Scheduler Duplication:**
- 170 individual systemd user timers (`job-auto-work-*.timer`) were active concurrently with the unified scheduler [`job-scheduler.py`](file:///home/super/Projects/NetVM/bin/job-scheduler.py), leading to excessive process spawns and duplicate dispatch checks.
4. **Lingering Scopes:**
- Previous browser watchdog relaunches occasionally killed only the Chromium child process while leaving transient systemd scopes and `netvm-cdp-relay.py` instances orphaned.
---
## 2. Remediations Applied
### A. Process & Scope Hygiene
- **Terminated Zombie Audio Streamer:**
```bash
kill 2661798 2661799
```
Immediately recovered ~1.0 off the system load average.
- **Cleared Orphaned Scopes:**
Stopped leftover transient scopes (`netvm-chrome-dev-1791158094`, `netvm-chrome-muse-1791217608`, `netvm-chrome-opm-1791217495`).
- **Hardened Watchdog Relaunch Logic:**
Updated [`bin/chromebox-watchdog.sh`](file:///home/super/Projects/NetVM/bin/chromebox-watchdog.sh) to explicitly stop existing user scopes matching `netvm-chrome-${PROFILE}-*` before starting a new transient unit.
### B. Chromium Headless Throttling & Memory Caps
Updated [`chrome-box`](file:///home/super/Projects/chrome-box/chrome-box) for both sandboxed and native (`--no-sandbox`) headless execution paths to include:
- `--disable-software-rasterizer` & `--disable-gpu-compositing` (disables CPU-heavy 3D software rasterization).
- `--mute-audio` (prevents unnecessary audio stream synthesis).
- `--disable-dev-shm-usage` (protects `/dev/shm` exhaustion).
- `--js-flags=--max-old-space-size=768` (bounds V8 heap consumption to 768MB per profile).
- Background timer and renderer throttling enabled (`--disable-renderer-backgrounding=false`, `--disable-background-timer-throttling=false`, `--disable-backgrounding-occluded-windows=false`).
### C. Timer Consolidation
- **Decommissioned 170 Redundant Timers:**
Archived all `job-auto-work-*.timer` and `.service` units to `~/.config/systemd/user/archive-auto-work-timers/` and disabled them.
- **Single Source of Truth:**
Retained [`job-scheduler.timer`](file:///home/super/.config/systemd/user/job-scheduler.timer) (`job-scheduler.py run` on `*:0/5`), which evaluates cron schedules in [`jobs/*.json`](file:///home/super/Projects/NetVM/jobs/) on a unified schedule. User timers dropped from **190 to 20**.
---
## 3. Verification & Fleet Health
Verification was conducted across all operational layers:
1. **Fleet Status:**
```bash
super fleet status
```
All 6 nodes (`muse`, `pip`, `646`, `opm`, `def`, `dev`) are **`● ACTIVE`** with sub-10ms CDP roundtrip latencies.
2. **Periodic Health Checks:**
```bash
tail -n 15 /tmp/agent-health.log
```
Consecutive runs of `agent-health.service` confirmed 100% OK status across all nodes.
3. **Scheduler Status:**
```bash
python3 bin/job-scheduler.py status
```
Unified scheduler runs cleanly every 5 minutes and updates `job-scheduler-state.json`.
4. **Log Retention:**
```bash
bin/retention-run-rotations.sh
```
Completed with return code `0`; archives passed gzip integrity and checksum validations.
5. **System Load:**
Load average dropped from **~5.8** to **~2.8–3.2** (~18% utilization across 16 logical cores).