Files
box/docs/FLEET-RESOURCE-OPTIMIZATION.md
T

4.3 KiB
Raw Blame History

Fleet Resource Optimization & Hygiene Runbook

Box is the main surface. All operator work goes through Box (box.muse-dev.online). The web UI, box CLI, and agents share the same API endpoints. No UI-only powers.

Date: 2026-10-05 ~17:40 UTC
Scope: Host bl (100.123.153.75), NetVM headless browser fleet, user systemd timers, and resource throttling.


1. Problem Summary & Root Causes

During operational monitoring, host CPU load average was observed elevated at 4.90–5.83. Investigation revealed three distinct contributors:

  1. Runaway Orphaned Audio Process:
    • PID 2661799 (parec -d gui_sink.monitor --format=s16le --rate=48000 --channels=2) was spinning in state R (runnable) at 100% of a CPU core.
    • It had accumulated 46,718 minutes (~778 hours) of continuous CPU time since September 3 after its downstream pipe reader terminated.
  2. SwiftShader Software WebGL Overhead:
    • Headless Chromium instances running without hardware acceleration were invoking SwiftShader (--use-angle=swiftshader-webgl) software rasterization loops on background tab animations.
  3. Timer & Scheduler Duplication:
    • 170 individual systemd user timers (job-auto-work-*.timer) were active concurrently with the unified scheduler job-scheduler.py, leading to excessive process spawns and duplicate dispatch checks.
  4. Lingering Scopes:
    • Previous browser watchdog relaunches occasionally killed only the Chromium child process while leaving transient systemd scopes and netvm-cdp-relay.py instances orphaned.

2. Remediations Applied

A. Process & Scope Hygiene

  • Terminated Zombie Audio Streamer:
    kill 2661798 2661799
    
    Immediately recovered ~1.0 off the system load average.
  • Cleared Orphaned Scopes: Stopped leftover transient scopes (netvm-chrome-dev-1791158094, netvm-chrome-muse-1791217608, netvm-chrome-opm-1791217495).
  • Hardened Watchdog Relaunch Logic: Updated bin/chromebox-watchdog.sh to explicitly stop existing user scopes matching netvm-chrome-${PROFILE}-* before starting a new transient unit.

B. Chromium Headless Throttling & Memory Caps

Updated chrome-box for both sandboxed and native (--no-sandbox) headless execution paths to include:

  • --disable-software-rasterizer & --disable-gpu-compositing (disables CPU-heavy 3D software rasterization).
  • --mute-audio (prevents unnecessary audio stream synthesis).
  • --disable-dev-shm-usage (protects /dev/shm exhaustion).
  • --js-flags=--max-old-space-size=768 (bounds V8 heap consumption to 768MB per profile).
  • Background timer and renderer throttling enabled (--disable-renderer-backgrounding=false, --disable-background-timer-throttling=false, --disable-backgrounding-occluded-windows=false).

C. Timer Consolidation

  • Decommissioned 170 Redundant Timers: Archived all job-auto-work-*.timer and .service units to ~/.config/systemd/user/archive-auto-work-timers/ and disabled them.
  • Single Source of Truth: Retained job-scheduler.timer (job-scheduler.py run on *:0/5), which evaluates cron schedules in jobs/*.json on a unified schedule. User timers dropped from 190 to 20.

3. Verification & Fleet Health

Verification was conducted across all operational layers:

  1. Fleet Status:
    super fleet status
    
    All 6 nodes (muse, pip, 646, opm, def, dev) are ● ACTIVE with sub-10ms CDP roundtrip latencies.
  2. Periodic Health Checks:
    tail -n 15 /tmp/agent-health.log
    
    Consecutive runs of agent-health.service confirmed 100% OK status across all nodes.
  3. Scheduler Status:
    python3 bin/job-scheduler.py status
    
    Unified scheduler runs cleanly every 5 minutes and updates job-scheduler-state.json.
  4. Log Retention:
    bin/retention-run-rotations.sh
    
    Completed with return code 0; archives passed gzip integrity and checksum validations.
  5. System Load: Load average dropped from ~5.8 to ~2.8–3.2 (~18% utilization across 16 logical cores).