Files
box/jobs/exec-sigkill-investigation-20261005.json
T

15 lines
1.6 KiB
JSON
Raw Normal View History

{
"agent": "646",
"chain_next": null,
"description": "Investigate exec-constrained.py SIGKILL/wedged-process cluster on bl (2026-10-05 03:37-05:40 UTC). Read-only forensics first.",
"dm_target": "646 tasks",
"name": "exec-sigkill-investigation-20261005",
"on_failure": "alert",
"prompt_template": "Investigate the exec-constrained.py SIGKILL/wedged-process cluster on bl (2026-10-05 03:37-05:40 UTC).\n\nJob ID: {job_id}\nTime: {datetime}\n\nBACKGROUND: Read-only deployment reconnaissance found 18 exec-server restarts on 2026-10-05. Notable: SIGKILLs at 03:37, 03:39, 03:49 UTC; wedged-process (Failed to kill) events at 04:19 and 05:40 UTC. The 07:50/09:50/12:51 UTC 502 windows had NO restarts (those were head-of-line blocking, now fixed by ThreadingHTTPServer deploy).\n\nSCOPE (read-only first):\n1. Pull journalctl for exec-constrained.service around 03:30-06:00 UTC 2026-10-05.\n2. Check for OOM-killer activity (dmesg, journal) around 03:37/03:39/03:49.\n3. Identify what sent SIGKILL (manual? OOM? systemd?).\n4. Investigate the 04:19/05:40 wedged processes: what child subprocesses were unkillable? Check for zombie/defunct processes.\n5. Review memory/cgroup pressure history if available.\n6. Correlate with the 18 restarts: are they all explained?\n\nDO NOT restart services or modify config. Read-only forensics.\n\nReport back with: (a) root cause hypothesis for SIGKILLs, (b) root cause for wedged processes, (c) recommended fix (needs user approval before any change).",
"sidechat": {
"create": false
},
"timeout": 600,
"schedule": "30 3 1 * *"
}