Files
box/docs/TIMER-STAGGER.md

4.3 KiB

Timer Stagger — Fleet Polling De-correlation

Box is the main surface. All operator work goes through Box (box.muse-dev.online). The web UI, box CLI, and agents share the same API endpoints. No UI-only powers.

Date: 2026-10-04 ~21:50 UTC Why: All 4 NetVM nodes egress via the same Warp IP (104.28.195.181). Synchronized timers = 4x correlated traffic bursts from one IP = correlated rate-limit risk. Staggering spreads load; combined with per-node rate limiting it de-correlates the fleet's egress signature.

Before

Three heaviest muse.ai-egress timers fired at the exact same second every 5 min (OnCalendar=*:0/5): job-heartbeat, self-main-loop, main-chat-watchdog. accounts-health + loop-remediator coincided every 15 min. All four wake-* timers fired hourly at :00. Three of four chromebox-watchdog timers were synchronized to the second.

After — stagger schedule

bl user timers (~/.config/systemd/user/, hand-managed)

Timer Old New Fires at
job-heartbeat *:0/5 *:0/5 (phase 0, unchanged) :00 :05 :10 …
self-main-loop *:0/5 *:1/5 :01 :06 :11 …
main-chat-watchdog *:0/5 *:2/5 :02 :07 :12 …
keepalive (box-managed) *:0/5 *:4/5 :04 :09 :14 …
accounts-health *:0/15 *:3/15 :03 :18 :33 :48
loop-remediator *:0/15 *:8/15 :08 :23 :38 :53
followup-sweeper *:* (@:00s) *:*:20 :20s of every minute
agent-health OnBootSec/OnUnitActiveSec 5m + RandomizedDelaySec=60s jittered ±60s
fleet-alert-check OnBootSec/OnUnitActiveSec 5m + RandomizedDelaySec=60s jittered ±60s
exec-watch OnUnitActiveSec 15m + RandomizedDelaySec=120s jittered ±120s

systemd natively supports offset calendar specs (*:1/5); no wrapper scripts needed.

Box-managed timers (/srv/box/timers.json → timer-sync.py --all)

Timer Old New
wake-muse *-*-* *:00:00 *-*-* *:05:00
wake-pip *-*-* *:00:00 *-*-* *:20:00
wake-646 *:0/30 *-*-* *:35:00
wake-opm *:0/30 *-*-* *:50:00
loop-health *:0/15 *:11/15
keepalive *:0/5 *:4/5

(wake-646/wake-opm were actually on *:0/30, not hourly — now spread 15 min apart across the hour like the others.)

System timers (/etc/systemd/system/, hand-managed)

Timer Change
chromebox-watchdog-{pip,646,muse,opm} + RandomizedDelaySec=30s (local checks only; de-correlates restart bursts)

Backups (pre-change)

  • bl user timers: ~/.config/systemd/user/bak-stagger-20261004/ (12 files)
  • system timers: /root/bak-stagger-20261004/
  • timers.json: /srv/box/timers.json.bak-stagger-20261004

Verification (2026-10-04 ~21:52 UTC)

systemctl --user list-timers and systemctl list-timers confirm the new phases: followup-sweeper @ :50:20, main-chat-watchdog @ :52:00, loop-remediator @ :53:00, accounts-health @ 22:03:00, keepalive @ :54:00, wake-muse @ 22:05, wake-pip @ 22:20, wake-646 @ 22:35, wake-opm @ 22:50. Watchdogs spread over ~7s via jitter.

Note: self-main-loop did one immediate catch-up run at 21:49:21 on reload (Persistent=true treating the 21:46 slot as missed under the new phase) — benign, one extra run.

Rollback

# bl user timers
cp ~/.config/systemd/user/bak-stagger-20261004/*.timer ~/.config/systemd/user/
systemctl --user daemon-reload
# system timers
sudo cp /root/bak-stagger-20261004/*.timer /etc/systemd/system/
sudo systemctl daemon-reload
# box timers
cp /srv/box/timers.json.bak-stagger-20261004 /srv/box/timers.json
/srv/box/bin/timer-sync.py --all

Not changed (deliberately)

  • box-http-health (*:0/15) / box-service-health (*:7/15) — already staggered.
  • siphon, cdp-relay-watchdog, shadow, backup-sync — unrelated cadences.
  • chat-state, response-harvester — inactive.
  • self-main-loop still touches all 4 browsers inside one run — staggering the timer doesn't de-correlate within the run. Follow-up: add inter-node sleeps inside self_main_loop.py itself.

Status

Live on bl + VM as of 2026-10-04 ~21:52 UTC. Doc uncommitted (review-then-commit); timer unit files live outside the NetVM repo (~/.config/systemd/user, /etc/systemd/system) — consider versioning them if drift recurs.