Files
multi-agent-mux/.agents/skills/multi-agent-mux-monitor/SKILL.md
T

12 KiB

name, description, version, author, license, platforms, environments, metadata
name description version author license platforms environments metadata
multi-agent-mux-monitor Run a long-lived reconciler that watches .mam/agent-sessions.yaml against the actual herdr/agent runtime state and reconciles them. Use when you want live visibility into which agent sessions are running, which are dead, which have stale YAML entries, and which have new session ids that haven't been recorded yet. Runs as a persistent loop (`reconcile.sh --subscribe`) that keeps going until it times out, idles out, or is interrupted. 1.0.0 godopu MIT
linux
macos
terminal
herdr
hermes
tags related_skills prereq_skills
agent
herdr
claude
antigravity
agy
monitor
observation
reconciliation
multi-agent-mux-create
multi-agent-mux-resume
multi-agent-mux-stop
multi-agent-mux-status
multi-agent-mux-create

Agent Sessions Monitor — Live Reconciliation

Companion skills: multi-agent-mux-create / multi-agent-mux-resume / multi-agent-mux-stop (mutators); this skill is the observer. Single source of truth: ./.mam/agent-sessions.yaml.

What this skill does

Run a long-lived reconciler (reconcile.sh --subscribe) that:

  1. Reacts to delegated-job events on the MQTT broker, and — whenever the broker is unreachable — falls back to polling every RECONCILE_POLL_INTERVAL (default 15s) the actual state of:
    • herdr agent list (which sessions are alive)
    • herdr agent get <session> (pane cmd, cwd)
    • ~/.claude/projects/<workspace-key>/*.jsonl mtime + first-line sessionId
    • ~/.gemini/antigravity-cli/cache/last_conversations.json (agy workspace → conversation mapping)
    • ~/.gemini/antigravity-cli/conversations/<uuid>.db mtime (agy)
  2. Compares the live state to agent-sessions.yaml
  3. Detects 4 classes of drift:
    • yaml-only terminated/archived/stopped: herdr dead, YAML says terminated, archived, or stopped → OK, left untouched (deliberate end states)
    • yaml-only running, herdr dead: YAML says running, herdr is gone → mark terminated with timestamp
    • herdr-only running, not in YAML: herdr session exists with <workspace>-creator-* naming but YAML doesn't know about it → register as a new entry
    • stale UUID: YAML has a UUID, but the on-disk artifact is gone → report it
  4. Emits a JSON drift record on stdout for every drift event when run with --emit-diff (note: the --subscribe broker-down fallback runs each pass for its YAML side-effects and discards the JSON — capture drift output with an explicit --once --emit-diff)
  5. Keeps running until one of its exit conditions fires: --timeout (wall-clock), --idle-timeout (no message received), or an interrupt from the operator.

When to use

  • You have multiple workspaces with herdr agent sessions and want a single source of truth
  • You suspect YAML drift after a host reboot / crash
  • You want a notification when a session id was just created (so you can record it before next restart)
  • You're running multi-day work and want to know "what's actually running right now"

When NOT to use

  • One-off interactive session — just check herdr agent list and read the YAML
  • A single, short session — overhead > benefit
  • You only need a point-in-time answer — use multi-agent-mux-status instead

Running the monitor

# Persistent monitor: runs until interrupted; polls if the broker is unreachable.
bash .agents/skills/multi-agent-mux-monitor/scripts/reconcile.sh --subscribe --idle-timeout 0

# Bounded run: exits after 5 min with no message, or 1 h wall-clock, whichever comes first.
bash .agents/skills/multi-agent-mux-monitor/scripts/reconcile.sh --subscribe --idle-timeout 300 --timeout 3600

Run it under whatever supervisor you already use (a dedicated herdr pane, nohup, or a background job). Nothing else needs to be running for the monitor to work — it reconciles YAML ↔ herdr ↔ disk on its own.

The herdr commands the script issues (herdr agent list, herdr agent get <session>) are real native herdr commands — do not substitute tmux-era names like herdr ls / herdr list-panes outside a shell that has sourced .agents/skills/lib.sh.

Helper script: reconcile.sh

This is the whole monitor — there is no separate driver. Each pass:

  1. Diffs YAML ↔ herdr ↔ disk artifacts
  2. Updates YAML if needed (only when changes are real, not on every poll — avoids spamming)
  3. Emits a JSON diff to stdout for the caller to consume
# Reconcile + auto-update YAML (atomic, flock-guarded). Emits JSON drift to stdout.
bash .agents/skills/multi-agent-mux-monitor/scripts/reconcile.sh --once --emit-diff

# Read-only: compute drift WITHOUT writing the YAML (use for "what's running?" checks).
bash .agents/skills/multi-agent-mux-monitor/scripts/reconcile.sh --once --emit-diff --dry-run

Flags: --once (single pass), --emit-diff (print JSON), --dry-run (P1-E — no mutation), --subscribe (push-based MQTT subscription monitoring). --subscribe sub-flags: --timeout N (exit after N seconds of wall-clock; 0 = no limit, default), --idle-timeout N (exit after N seconds with no message; default 3600, 0 = never idle-out). On a broker connection failure (connect error or non-zero CONNACK), --subscribe falls back to a polling loop that re-runs --once --emit-diff every RECONCILE_POLL_INTERVAL (default 15) seconds until --timeout. Terminal-event YAML updates are written through lib.sh::atomic_dump_yaml (flock + schema-validate + .bak). There are no --workspace / --agent flags; the emitted JSON drifts[] is the caller's to consume.

Drift classes (what the script handles)

Status Enum

The status field MUST be one of the following exact strings: running, stopped, terminated, archived. The last_visible_status is a free-form human-readable status string (e.g. verification-cycle states: unverified, pinned, resume_verified, or a failure detail string) and is NOT constrained to this enum. Any unstructured comments or reasons for the status change should be placed in last_visible_note or termination_mode.

A. herdr dead, YAML says running → auto-terminate

YAML:  status=running, pane.pid=201132, cmd=claude
herdr:  no session
       → set status=terminated, terminated_at=<now>, termination_mode=auto-detected
       → report: "lab-landing-page-creator-claude: herdr gone (was pane 201132, cmd claude). Marked terminated."

Skip-set: the auto-terminate only fires for sessions whose status is running. Rows already in a deliberate end state — terminated, archived, or stopped (set by multi-agent-mux-stop) — are left untouched. This is critical: a stopped row keeps its resumable: true and captured *_session_id_own, so the monitor must not overwrite it with terminated ("auto-detected") when its herdr is (expectedly) gone.

B. herdr alive, not in YAML → auto-register

herdr:  session=lab-paper-pdf2md-creator-agy, pid=...,
       cmd=agy, cwd=$WORKSPACE_ROOT/paper-pdf2md
YAML:  no such session
       → register as new entry: status=running, last_visible_status=running, last_visible_note=auto-registered
       → report: "lab-paper-pdf2md-creator-agy: herdr found but not in YAML. Auto-registered."

C. Session ID Discovery & Confirmation

  • C0. Assigned ID Confirmation: For sessions created with an auto-assigned UUID (session_id_source: assigned, session_id_verified: false), the monitor verifies that the transcript file .jsonl has materialized on disk. Once verified, it promotes session_id_verified: true and updates last_visible_status: pinned.
  • C. New session id materializes (unassigned/legacy):
    YAML:  claude_session_id_own=null (placeholder)
    disk:  ~/.claude/projects/.../b3a7...c2f.jsonl exists, mtime=now,
           first line sessionId=b3a7...c2f
           → update claude_session_id_own=b3a7...c2f
           → report: "lab-landing-page-creator-claude: session id materialized b3a7...c2f"
    
  • C-ambiguous. Multiple candidates detected: If multiple candidate transcripts match an unassigned session, the monitor avoids random pinning, reports C-ambiguous, and sets last_visible_status: "ambiguous: N candidates".

D. Stale UUID (artifact gone)

YAML:  agent_identities.claude.session_id=87dc548e-...
disk:  ~/.claude/projects/.../87dc548e-...jsonl: missing
       → report it, but DO NOT delete from YAML
       (the user may have moved the file or the disk may be temporarily unavailable;
        only `--purge-conversation` should remove the id)

Pitfalls

  • Don't expect --once to stay alive — it does a single pass and exits. Use --subscribe for continuous monitoring.
  • --idle-timeout defaults to 3600s — a monitor meant to run indefinitely needs --idle-timeout 0 explicitly, or it will quietly exit after an hour of broker silence.
  • The poll interval is a defaultRECONCILE_POLL_INTERVAL (15s) is what the broker-down fallback uses. A workspace with 5+ agent sessions can bump it to reduce noise.
  • Coalesce repeated drifts — the same drift re-appears on every pass until it is resolved. A caller that acts on drifts[] should compare against the previous pass and act only on new entries; the script does not deduplicate for you.
  • Don't fight the user's explicit action — if multi-agent-mux-stop is mid-flight and the monitor sees the same session in two states within 5s, prefer the user's most recent action. The monitor should not auto-revert a fresh terminated to running because of a stale herdr has-session check.
  • The monitor should never modify the conversation artifacts (jsonl, db) — only the YAML. If you see a stale UUID, report it but don't delete the file.
  • TUI capture-pane is expensive — only capture when you need to update last_visible_status, not every poll.

Supervising-agent runbook

If an agent drives the monitor rather than an operator watching it directly, this is the behavior spec:

# agent-sessions monitor

## Loop

1. Read agent-sessions.yaml
2. Bash: `bash .agents/skills/multi-agent-mux-monitor/scripts/reconcile.sh --once --emit-diff`
3. Parse the JSON diff from stdout
4. If `drifts` is non-empty, report each *new* drift to the operator
5. Bash: `sleep 30`, then repeat

## Stop condition

Stop when the operator says to stop, or when the surrounding job's timeout fires.

## Drift responses

- A. herdr dead + YAML running: auto-terminate YAML, report
- B. herdr alive not in YAML: auto-register, report
- C. New session id from *.jsonl: update YAML, report
- D. Stale UUID: report only, no YAML change

## Hard rules

- Do NOT modify conversation artifacts (jsonl, db, brain/)
- Do NOT spawn/delete herdr sessions — that's the create/delete skills' job
- Do NOT call multi-agent-mux-create or multi-agent-mux-stop — only the user initiates those
- Do NOT call `git commit` / `git push`

Security: --subscribe on Public Brokers

When using --subscribe with the default PoC public broker (broker.hivemq.com:1883), be aware that:

  1. Wildcard subscription means anyone can publish events to your job topics.
  2. Auto-kill on terminal events means a spoofed completed or error event from a third party can terminate your agent session.
  3. Mitigation: Use --subscribe only on private TLS-enabled brokers (production mode). For PoC, prefer polling-based monitor (--once or no --subscribe) which reads YAML/herdr state directly without MQTT.
  4. HMAC verification: Events are now verified via verify_hmac() in mqtt_common.py (see FW-05). Ensure auth_token is set for each job to enable signature validation — unauthenticated events will be dropped.

Verification (one-shot)

# Run reconcile once and inspect output
bash .agents/skills/multi-agent-mux-monitor/scripts/reconcile.sh --emit-diff --once \
  | python3 -m json.tool
  • multi-agent-mux-status — read-only snapshot when you don't need a running loop
  • multi-agent-mux-delegate-job — the MQTT job channel whose events --subscribe listens to