12 KiB
name, description, version, author, license, platforms, environments, metadata
| name | description | version | author | license | platforms | environments | metadata | ||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| multi-agent-mux-monitor | Run a long-lived reconciler that watches .mam/agent-sessions.yaml against the actual herdr/agent runtime state and reconciles them. Use when you want live visibility into which agent sessions are running, which are dead, which have stale YAML entries, and which have new session ids that haven't been recorded yet. Runs as a persistent loop (`reconcile.sh --subscribe`) that keeps going until it times out, idles out, or is interrupted. | 3.0.0 | godopu | MIT |
|
|
|
Agent Sessions Monitor — Live Reconciliation
Companion skills:
multi-agent-mux-create/multi-agent-mux-resume/multi-agent-mux-stop(mutators); this skill is the observer. Single source of truth:./.mam/agent-sessions.yaml.
What this skill does
Run a long-lived reconciler (reconcile.sh --subscribe) that:
- Reacts to delegated-job events on the MQTT broker, and — whenever the broker is
unreachable — falls back to polling every
RECONCILE_POLL_INTERVAL(default 15s) the actual state of:herdr agent list(which sessions are alive)herdr agent get <session>(pane cmd, cwd)~/.claude/projects/<workspace-key>/*.jsonlmtime + first-line sessionId~/.gemini/antigravity-cli/cache/last_conversations.json(agy workspace → conversation mapping)~/.gemini/antigravity-cli/conversations/<uuid>.dbmtime (agy)
- Compares the live state to
agent-sessions.yaml - Detects 4 classes of drift:
- yaml-only terminated/archived/stopped: herdr dead, YAML says
terminated,archived, orstopped→ OK, left untouched (deliberate end states) - yaml-only running, herdr dead: YAML says
running, herdr is gone → markterminatedwith timestamp - herdr-only running, not in YAML: herdr session exists with
<workspace>-creator-*naming but YAML doesn't know about it → register as a new entry - stale UUID: YAML has a UUID, but the on-disk artifact is gone → report it
- yaml-only terminated/archived/stopped: herdr dead, YAML says
- Emits a JSON drift record on stdout for every drift event when run with
--emit-diff(note: the--subscribebroker-down fallback runs each pass for its YAML side-effects and discards the JSON — capture drift output with an explicit--once --emit-diff) - Keeps running until one of its exit conditions fires:
--timeout(wall-clock),--idle-timeout(no message received), or an interrupt from the operator.
When to use
- You have multiple workspaces with herdr agent sessions and want a single source of truth
- You suspect YAML drift after a host reboot / crash
- You want a notification when a session id was just created (so you can record it before next restart)
- You're running multi-day work and want to know "what's actually running right now"
When NOT to use
- One-off interactive session — just check
herdr agent listand read the YAML - A single, short session — overhead > benefit
- You only need a point-in-time answer — use
multi-agent-mux-statusinstead
Running the monitor
# Persistent monitor: runs until interrupted; polls if the broker is unreachable.
bash .agents/skills/multi-agent-mux-monitor/scripts/reconcile.sh --subscribe --idle-timeout 0
# Bounded run: exits after 5 min with no message, or 1 h wall-clock, whichever comes first.
bash .agents/skills/multi-agent-mux-monitor/scripts/reconcile.sh --subscribe --idle-timeout 300 --timeout 3600
Run it under whatever supervisor you already use (a dedicated herdr pane, nohup,
or a background job). Nothing else needs to be running for the monitor to work —
it reconciles YAML ↔ herdr ↔ disk on its own.
The herdr commands the script issues (herdr agent list, herdr agent get <session>)
are real native herdr commands — do not substitute tmux-era names like herdr ls /
herdr list-panes outside a shell that has sourced .agents/skills/lib.sh.
Helper script: reconcile.sh
This is the whole monitor — there is no separate driver. Each pass:
- Diffs YAML ↔ herdr ↔ disk artifacts
- Updates YAML if needed (only when changes are real, not on every poll — avoids spamming)
- Emits a JSON diff to stdout for the caller to consume
# Reconcile + auto-update YAML (atomic, flock-guarded). Emits JSON drift to stdout.
bash .agents/skills/multi-agent-mux-monitor/scripts/reconcile.sh --once --emit-diff
# Read-only: compute drift WITHOUT writing the YAML (use for "what's running?" checks).
bash .agents/skills/multi-agent-mux-monitor/scripts/reconcile.sh --once --emit-diff --dry-run
Flags: --once (single pass), --emit-diff (print JSON), --dry-run (P1-E — no mutation), --subscribe (push-based MQTT subscription monitoring). --subscribe sub-flags: --timeout N (exit after N seconds of wall-clock; 0 = no limit, default), --idle-timeout N (exit after N seconds with no message; default 3600, 0 = never idle-out). On a broker connection failure (connect error or non-zero CONNACK), --subscribe falls back to a polling loop that re-runs --once --emit-diff every RECONCILE_POLL_INTERVAL (default 15) seconds until --timeout. Terminal-event YAML updates are written through lib.sh::atomic_dump_yaml (flock + schema-validate + .bak). There are no --workspace / --agent flags; the emitted JSON drifts[] is the caller's to consume.
Drift classes (what the script handles)
Status Enum
The status field MUST be one of the following exact strings: running, stopped, terminated, archived.
The last_visible_status is a free-form human-readable status string (e.g. verification-cycle states: unverified, pinned, resume_verified, or a failure detail string) and is NOT constrained to this enum.
Any unstructured comments or reasons for the status change should be placed in last_visible_note or termination_mode.
A. herdr dead, YAML says running → auto-terminate
YAML: status=running, pane.pid=201132, cmd=claude
herdr: no session
→ set status=terminated, terminated_at=<now>, termination_mode=auto-detected
→ report: "lab-landing-page-creator-claude: herdr gone (was pane 201132, cmd claude). Marked terminated."
Skip-set: the auto-terminate only fires for sessions whose status is running.
Rows already in a deliberate end state — terminated, archived, or stopped
(set by multi-agent-mux-stop) — are
left untouched. This is critical: a stopped row keeps its resumable: true and
captured *_session_id_own, so the monitor must not overwrite it with
terminated ("auto-detected") when its herdr is (expectedly) gone.
B. herdr alive, not in YAML → auto-register
herdr: session=lab-paper-pdf2md-creator-agy, pid=...,
cmd=agy, cwd=$WORKSPACE_ROOT/paper-pdf2md
YAML: no such session
→ register as new entry: status=running, last_visible_status=running, last_visible_note=auto-registered
→ report: "lab-paper-pdf2md-creator-agy: herdr found but not in YAML. Auto-registered."
C. Session ID Discovery & Confirmation
- C0. Assigned ID Confirmation: For sessions created with an auto-assigned UUID (
session_id_source: assigned,session_id_verified: false), the monitor verifies that the transcript file.jsonlhas materialized on disk. Once verified, it promotessession_id_verified: trueand updateslast_visible_status: pinned. - C. New session id materializes (unassigned/legacy):
YAML: claude_session_id_own=null (placeholder) disk: ~/.claude/projects/.../b3a7...c2f.jsonl exists, mtime=now, first line sessionId=b3a7...c2f → update claude_session_id_own=b3a7...c2f → report: "lab-landing-page-creator-claude: session id materialized b3a7...c2f" - C-ambiguous. Multiple candidates detected: If multiple candidate transcripts match an unassigned session, the monitor avoids random pinning, reports
C-ambiguous, and setslast_visible_status: "ambiguous: N candidates".
Pitfalls
- Don't expect
--onceto stay alive — it does a single pass and exits. Use--subscribefor continuous monitoring. --idle-timeoutdefaults to 3600s — a monitor meant to run indefinitely needs--idle-timeout 0explicitly, or it will quietly exit after an hour of broker silence.- The poll interval is a default —
RECONCILE_POLL_INTERVAL(15s) is what the broker-down fallback uses. A workspace with 5+ agent sessions can bump it to reduce noise. - Coalesce repeated drifts — the same drift re-appears on every pass until it is resolved. A caller that acts on
drifts[]should compare against the previous pass and act only on new entries; the script does not deduplicate for you. - Don't fight the user's explicit action — if
multi-agent-mux-stopis mid-flight and the monitor sees the same session in two states within 5s, prefer the user's most recent action. The monitor should not auto-revert a freshterminatedtorunningbecause of a staleherdr has-sessioncheck. - The monitor should never modify the conversation artifacts (jsonl, db) — only the YAML. If you see a stale UUID, report it but don't delete the file.
- TUI capture-pane is expensive — only capture when you need to update
last_visible_status, not every poll.
Supervising-agent runbook
If an agent drives the monitor rather than an operator watching it directly, this is the behavior spec:
# agent-sessions monitor
## Loop
1. Read agent-sessions.yaml
2. Bash: `bash .agents/skills/multi-agent-mux-monitor/scripts/reconcile.sh --once --emit-diff`
3. Parse the JSON diff from stdout
4. If `drifts` is non-empty, report each *new* drift to the operator
5. Bash: `sleep 30`, then repeat
## Stop condition
Stop when the operator says to stop, or when the surrounding job's timeout fires.
## Drift responses
- A. herdr dead + YAML running: auto-terminate YAML, report
- B. herdr alive not in YAML: auto-register, report
- C. New session id from *.jsonl: update YAML, report
- D. Stale UUID: report only, no YAML change
## Hard rules
- Do NOT modify conversation artifacts (jsonl, db, brain/)
- Do NOT spawn/delete herdr sessions — that's the create/delete skills' job
- Do NOT call multi-agent-mux-create or multi-agent-mux-stop — only the user initiates those
- Do NOT call `git commit` / `git push`
Security: --subscribe on Public Brokers
When using --subscribe with the default PoC public broker
(broker.hivemq.com:1883), be aware that:
- Wildcard subscription means anyone can publish events to your job topics.
- Auto-kill on terminal events means a spoofed
completedorerrorevent from a third party can terminate your agent session. - Mitigation: Use
--subscribeonly on private TLS-enabled brokers (production mode). For PoC, prefer polling-based monitor (--onceor no--subscribe) which reads YAML/herdr state directly without MQTT. - HMAC verification: Events are now verified via
verify_hmac()inmqtt_common.py(see FW-05). Ensureauth_tokenis set for each job to enable signature validation — unauthenticated events will be dropped.
Verification (one-shot)
# Run reconcile once and inspect output
bash .agents/skills/multi-agent-mux-monitor/scripts/reconcile.sh --emit-diff --once \
| python3 -m json.tool
Related skills
multi-agent-mux-status— read-only snapshot when you don't need a running loopmulti-agent-mux-delegate-job— the MQTT job channel whose events--subscribelistens to