Files
box/docs/JOB-SPEC.md
T

7.8 KiB

JOB-SPEC.md: Hosted Job Scheduler and Distributor

Box is the main surface. All operator work goes through Box (box.muse-dev.online). The web UI, box CLI, and agents share the same API endpoints. No UI-only powers.

Overview

A hosted system on bl that processes and distributes jobs to agents via DM. Cron jobs and system scripts inject prompts/jobs; agents execute them. Loops run as loops on the server, not in agent heads.

Principle: Limit agency to get smarter. The server decides what and when; agents decide how. Deterministic, auditable, debuggable.

Architecture

┌──────────────────────────────────────────────────┐
│ Hosted on bl (systemd timers + scripts)          │
│                                                  │
│  ┌────────────┐  ┌─────────────┐                 │
│  │ Scheduler  │→ │ Dispatcher  │→ dm.py send     │
│  │ (cron)     │  │ (render+log)│                 │
│  └────────────┘  └─────────────┘                 │
│        ↑                ↓                        │
│        │         ┌─────────────┐                  │
│        └─────────│ Collector   │← DM [RESULT]     │
│                  │ (log+chain) │                  │
│                  └─────────────┘                  │
└──────────────────────────────────────────────────┘

Components

1. Job Definition (YAML)

Location: /home/super/Projects/NetVM/jobs/<name>.yaml

name: board-watch
description: "Check board for new posts every 5 minutes"
schedule: "*/5 * * * *"  # cron format
agent: muse              # which agent executes
sidechat:
  create: false          # use main chat
  # OR:
  # create: true
  # name_template: "job-{name}-{date}"
  # reuse_pattern: "job-{name}-*"  # for chaining
prompt_template: |
  Check the board for posts since {last_run}.
  Summarize new items in 3 bullet points.
  Reply with [RESULT {job_id}] followed by your summary.
timeout: 300             # seconds before marking failed
on_failure: retry       # retry | alert | ignore
chain_next: null        # job to trigger after success

Fields:

  • name: Unique job identifier (used in logs, sidechat names)
  • schedule: Cron expression (systemd timer or cron)
  • agent: Target agent (muse, pip, 646, opm)
  • sidechat: Side chat configuration (see below)
  • prompt_template: Jinja-style template with {variables}
  • timeout: Max seconds to wait for result
  • on_failure: What to do on timeout/failure
  • chain_next: Next job to trigger (for chains)

2. Scheduler

Uses systemd timers (preferred) or cron. Each job gets a timer unit.

Systemd timer example:

# /etc/systemd/user/job-board-watch.timer
[Unit]
Description=Run board-watch job every 5 minutes

[Timer]
OnCalendar=*:0/5
Persistent=true

[Install]
WantedBy=timers.target

Service:

# /etc/systemd/user/job-board-watch.service
[Unit]
Description=Dispatch board-watch job

[Service]
Type=oneshot
ExecStart=/home/super/Projects/NetVM/bin/job-dispatch.sh board-watch

3. Dispatcher (bin/job-dispatch.sh)

Responsibilities:

  1. Load job YAML
  2. Render prompt template with variables ({job_id}, {last_run}, {date}, etc.)
  3. Create side chat if specified
  4. Send DM via dm.py:
    • Format: [JOB {job_id}] {rendered_prompt}
    • Use --raw if prompt is pre-signed
  5. Log to job-log.jsonl: {job_id, job_name, agent, sidechat, sent_at}
  6. If chain_next, schedule the next job (or trigger immediately on result)

Job ID format: {name}-{YYYYMMDD-HHMMSS}-{short_uuid} Example: board-watch-20261004-023000-a1b2c3d4

4. Agent Job Handler (Convention)

Agents MUST recognize job DMs and respond in format.

Job DM format:

[JOB board-watch-20261004-023000-a1b2c3d4] Check the board for posts
since 2026-10-04T02:25:00Z. Summarize new items in 3 bullet points.
Reply with [RESULT board-watch-20261004-023000-a1b2c3d4] followed by
your summary.

Agent responsibilities:

  1. Parse [JOB {job_id}] from DM
  2. Execute the prompt
  3. If sidechat specified, work in that side chat
  4. Reply via DM with: [RESULT {job_id}] {result_text}
  5. If unable, reply: [RESULT {job_id}] FAILED: {reason}

Teaching: New agents get JOB-HANDLER.md with examples. The format is simple enough to learn from 2-3 examples.

5. Collector

Watches for [RESULT {job_id}] in DM logs or via dm.py log.

Responsibilities:

  1. Parse result DMs
  2. Log to job-log.jsonl: {job_id, completed_at, success, result_preview}
  3. If chain_next specified and result was success, trigger next job
  4. On timeout (no result within timeout), mark failed, apply on_failure

Timeout handling: A background sweeper checks for jobs with sent_at older than timeout and no result. Marks them failed.

Side Chat Integration

Job → Side Chat Mapping

Jobs can specify side chat behavior:

Option A: No side chat (use main chat)

sidechat:
  create: false

Option B: Create new side chat per run

sidechat:
  create: true
  name_template: "job-{name}-{date}"  # e.g., "job-board-watch-20261004"

Option C: Reuse side chat (for chains)

sidechat:
  create: false
  reuse_pattern: "job-{name}-*"  # find most recent
  # OR:
  # reuse_name: "{prev_job_sidechat}"  # from chain

Side Chat as Job Workspace

When create: true:

  1. Dispatcher calls dm.py sidechat create (via muse-chat-api.py)
  2. Gets the side chat name/ID
  3. Includes it in the job DM: "Work in side chat 'job-board-watch-20261004'"
  4. Logs the mapping: {job_id → sidechat_name}
  5. Agent does all work in that side chat (full context, isolated)

Benefits:

  • Each job run has isolated context
  • The side chat IS the audit log
  • Chains: Job B can continue in Job A's side chat
  • No cross-talk between concurrent jobs

Logging

job-log.jsonl (on bl)

Append-only, one JSON per line:

{"ts": "2026-10-04T02:30:00Z", "type": "job_sent", "job_id": "...", "job_name": "board-watch", "agent": "muse", "sidechat": null}
{"ts": "2026-10-04T02:32:15Z", "type": "job_result", "job_id": "...", "success": true, "duration_s": 135}
{"ts": "2026-10-04T02:35:00Z", "type": "job_timeout", "job_id": "...", "job_name": "board-watch"}

sidechat-log.jsonl (on bl)

From sidechat_manager.py:

{"ts": "...", "agent": "muse", "op": "create", "details": "job-board-watch-20261004"}
{"ts": "...", "agent": "muse", "op": "navigate", "details": "main"}

Rate Limiting

Use the shared bin/rate_limiter.py module:

from rate_limiter import rate_limit_wait
rate_limit_wait(agent)  # blocks if too frequent

Default: 1 DM per 3s per agent, burst 5, max 20/min. Prevents accidental spam if a job misfires in a loop.

Security

  • Job definitions are in git (auditable)
  • Only operators can create/edit jobs (file permissions)
  • Agents cannot create jobs (they execute, not schedule)
  • DMs are logged (dm-log.jsonl) for audit
  • Side chats are per-job, not shared across trust boundaries

Future Expansions

  1. Conditional jobs: Run Job B only if Job A succeeds with specific output
  2. Parallel jobs: Fan-out to multiple agents, collect all results
  3. Human approval: Certain jobs require human sign-off before dispatch
  4. Web UI: View job status, logs, side chats from box.muse-dev.online
  5. Agent-proposed jobs: Agents can suggest jobs via [PROPOSE] DM, human approves

Status

Spec v0.1. Owner: operator-main. Sanctioned by the human 2026-10-04. Implementation: dispatcher script + example jobs pending.