4.3 KiB
Timer Stagger — Fleet Polling De-correlation
Box is the main surface. All operator work goes through Box (box.muse-dev.online). The web UI,
boxCLI, and agents share the same API endpoints. No UI-only powers.
Date: 2026-10-04 ~21:50 UTC Why: All 4 NetVM nodes egress via the same Warp IP (104.28.195.181). Synchronized timers = 4x correlated traffic bursts from one IP = correlated rate-limit risk. Staggering spreads load; combined with per-node rate limiting it de-correlates the fleet's egress signature.
Before
Three heaviest muse.ai-egress timers fired at the exact same second every 5 min
(OnCalendar=*:0/5): job-heartbeat, self-main-loop, main-chat-watchdog.
accounts-health + loop-remediator coincided every 15 min. All four wake-*
timers fired hourly at :00. Three of four chromebox-watchdog timers were
synchronized to the second.
After — stagger schedule
bl user timers (~/.config/systemd/user/, hand-managed)
| Timer | Old | New | Fires at |
|---|---|---|---|
| job-heartbeat | *:0/5 |
*:0/5 (phase 0, unchanged) |
:00 :05 :10 … |
| self-main-loop | *:0/5 |
*:1/5 |
:01 :06 :11 … |
| main-chat-watchdog | *:0/5 |
*:2/5 |
:02 :07 :12 … |
| keepalive (box-managed) | *:0/5 |
*:4/5 |
:04 :09 :14 … |
| accounts-health | *:0/15 |
*:3/15 |
:03 :18 :33 :48 |
| loop-remediator | *:0/15 |
*:8/15 |
:08 :23 :38 :53 |
| followup-sweeper | *:* (@:00s) |
*:*:20 |
:20s of every minute |
| agent-health | OnBootSec/OnUnitActiveSec 5m | + RandomizedDelaySec=60s |
jittered ±60s |
| fleet-alert-check | OnBootSec/OnUnitActiveSec 5m | + RandomizedDelaySec=60s |
jittered ±60s |
| exec-watch | OnUnitActiveSec 15m | + RandomizedDelaySec=120s |
jittered ±120s |
systemd natively supports offset calendar specs (*:1/5); no wrapper scripts needed.
Box-managed timers (/srv/box/timers.json → timer-sync.py --all)
| Timer | Old | New |
|---|---|---|
| wake-muse | *-*-* *:00:00 |
*-*-* *:05:00 |
| wake-pip | *-*-* *:00:00 |
*-*-* *:20:00 |
| wake-646 | *:0/30 |
*-*-* *:35:00 |
| wake-opm | *:0/30 |
*-*-* *:50:00 |
| loop-health | *:0/15 |
*:11/15 |
| keepalive | *:0/5 |
*:4/5 |
(wake-646/wake-opm were actually on *:0/30, not hourly — now spread 15 min apart
across the hour like the others.)
System timers (/etc/systemd/system/, hand-managed)
| Timer | Change |
|---|---|
| chromebox-watchdog-{pip,646,muse,opm} | + RandomizedDelaySec=30s (local checks only; de-correlates restart bursts) |
Backups (pre-change)
- bl user timers:
~/.config/systemd/user/bak-stagger-20261004/(12 files) - system timers:
/root/bak-stagger-20261004/ - timers.json:
/srv/box/timers.json.bak-stagger-20261004
Verification (2026-10-04 ~21:52 UTC)
systemctl --user list-timers and systemctl list-timers confirm the new phases:
followup-sweeper @ :50:20, main-chat-watchdog @ :52:00, loop-remediator @ :53:00,
accounts-health @ 22:03:00, keepalive @ :54:00, wake-muse @ 22:05, wake-pip @ 22:20,
wake-646 @ 22:35, wake-opm @ 22:50. Watchdogs spread over ~7s via jitter.
Note: self-main-loop did one immediate catch-up run at 21:49:21 on reload
(Persistent=true treating the 21:46 slot as missed under the new phase) — benign,
one extra run.
Rollback
# bl user timers
cp ~/.config/systemd/user/bak-stagger-20261004/*.timer ~/.config/systemd/user/
systemctl --user daemon-reload
# system timers
sudo cp /root/bak-stagger-20261004/*.timer /etc/systemd/system/
sudo systemctl daemon-reload
# box timers
cp /srv/box/timers.json.bak-stagger-20261004 /srv/box/timers.json
/srv/box/bin/timer-sync.py --all
Not changed (deliberately)
box-http-health(*:0/15) /box-service-health(*:7/15) — already staggered.siphon,cdp-relay-watchdog,shadow,backup-sync— unrelated cadences.chat-state,response-harvester— inactive.self-main-loopstill touches all 4 browsers inside one run — staggering the timer doesn't de-correlate within the run. Follow-up: add inter-node sleeps insideself_main_loop.pyitself.
Status
Live on bl + VM as of 2026-10-04 ~21:52 UTC. Doc uncommitted (review-then-commit);
timer unit files live outside the NetVM repo (~/.config/systemd/user,
/etc/systemd/system) — consider versioning them if drift recurs.