4.3 KiB
4.3 KiB
Fleet Resource Optimization & Hygiene Runbook
Box is the main surface. All operator work goes through Box (box.muse-dev.online). The web UI,
boxCLI, and agents share the same API endpoints. No UI-only powers.
Date: 2026-10-05 ~17:40 UTC
Scope: Host bl (100.123.153.75), NetVM headless browser fleet, user systemd timers, and resource throttling.
1. Problem Summary & Root Causes
During operational monitoring, host CPU load average was observed elevated at 4.90–5.83. Investigation revealed three distinct contributors:
- Runaway Orphaned Audio Process:
PID 2661799(parec -d gui_sink.monitor --format=s16le --rate=48000 --channels=2) was spinning in stateR(runnable) at 100% of a CPU core.- It had accumulated 46,718 minutes (~778 hours) of continuous CPU time since September 3 after its downstream pipe reader terminated.
- SwiftShader Software WebGL Overhead:
- Headless Chromium instances running without hardware acceleration were invoking SwiftShader (
--use-angle=swiftshader-webgl) software rasterization loops on background tab animations.
- Headless Chromium instances running without hardware acceleration were invoking SwiftShader (
- Timer & Scheduler Duplication:
- 170 individual systemd user timers (
job-auto-work-*.timer) were active concurrently with the unified schedulerjob-scheduler.py, leading to excessive process spawns and duplicate dispatch checks.
- 170 individual systemd user timers (
- Lingering Scopes:
- Previous browser watchdog relaunches occasionally killed only the Chromium child process while leaving transient systemd scopes and
netvm-cdp-relay.pyinstances orphaned.
- Previous browser watchdog relaunches occasionally killed only the Chromium child process while leaving transient systemd scopes and
2. Remediations Applied
A. Process & Scope Hygiene
- Terminated Zombie Audio Streamer:
Immediately recovered ~1.0 off the system load average.
kill 2661798 2661799 - Cleared Orphaned Scopes:
Stopped leftover transient scopes (
netvm-chrome-dev-1791158094,netvm-chrome-muse-1791217608,netvm-chrome-opm-1791217495). - Hardened Watchdog Relaunch Logic:
Updated
bin/chromebox-watchdog.shto explicitly stop existing user scopes matchingnetvm-chrome-${PROFILE}-*before starting a new transient unit.
B. Chromium Headless Throttling & Memory Caps
Updated chrome-box for both sandboxed and native (--no-sandbox) headless execution paths to include:
--disable-software-rasterizer&--disable-gpu-compositing(disables CPU-heavy 3D software rasterization).--mute-audio(prevents unnecessary audio stream synthesis).--disable-dev-shm-usage(protects/dev/shmexhaustion).--js-flags=--max-old-space-size=768(bounds V8 heap consumption to 768MB per profile).- Background timer and renderer throttling enabled (
--disable-renderer-backgrounding=false,--disable-background-timer-throttling=false,--disable-backgrounding-occluded-windows=false).
C. Timer Consolidation
- Decommissioned 170 Redundant Timers:
Archived all
job-auto-work-*.timerand.serviceunits to~/.config/systemd/user/archive-auto-work-timers/and disabled them. - Single Source of Truth:
Retained
job-scheduler.timer(job-scheduler.py runon*:0/5), which evaluates cron schedules injobs/*.jsonon a unified schedule. User timers dropped from 190 to 20.
3. Verification & Fleet Health
Verification was conducted across all operational layers:
- Fleet Status:
All 6 nodes (
super fleet statusmuse,pip,646,opm,def,dev) are● ACTIVEwith sub-10ms CDP roundtrip latencies. - Periodic Health Checks:
Consecutive runs of
tail -n 15 /tmp/agent-health.logagent-health.serviceconfirmed 100% OK status across all nodes. - Scheduler Status:
Unified scheduler runs cleanly every 5 minutes and updates
python3 bin/job-scheduler.py statusjob-scheduler-state.json. - Log Retention:
Completed with return code
bin/retention-run-rotations.sh0; archives passed gzip integrity and checksum validations. - System Load: Load average dropped from ~5.8 to ~2.8–3.2 (~18% utilization across 16 logical cores).