# Test Infrastructure Specification ## Test Philosophy We adopt an **opaque-box, requirement-driven** testing philosophy for the tmux-to-herdr migration scripts. This approach ensures that the test suite validates external behaviors, input/output contracts, and side-effects rather than asserting internal code structure or layout. The scripts are treated as black boxes that: - Accept CLI arguments and environment variables. - Query/interact with the `herdr` daemon through the `_herdr` shim (using mock executable interception). - Perform state mutations inside `.mam/agent-sessions.yaml` and `.mam/agent-sessions.db`. - Interact with background agent runners (`claude`, `agy`, `hermes`, `cline`). This guarantees that our test assertions remain stable even if the script implementation details are refactored, as long as the functional requirements are met. ## Feature Inventory The test suite is structured around five core features, mapping out verification checks across Tiers 1, 2, and 3: | Feature | Tier 1 (Unit Checks) | Tier 2 (Component Checks) | Tier 3 (Integration Checks) | |---|---|---|---| | **Create Session** | - `derive_session_name` slug generation checks
- Workspace-to-slug character translation
- Invalid workspace path filtering
- Role parameter sanity validations
- Session override string generation | - Verification of state serialization to YAML schema
- Isolation home directory structure validation
- SQLite DB connection verification
- Concurrency check for database registration lock
- Database schema validation on write | - Spawn session execution with mock `herdr` and mock agent
- TUI readiness wait check
- Cleanup trap execution on crash
- Argument validation logic verification
- Isolation directory creation checks | | **Resume Session** | - Workspace UUID resolution order unit tests
- CLI session ID parser validations
- Check prioritization (yaml file -> disk scan -> cache)
- Workspace path boundary check
- Empty UUID handling logic | - Configuration restore verification
- Environment overrides assertion
- Integrity check on retrieved SQLite metadata
- Validation of session ownership verification
- Config parsing for resume options | - Run `resume_session.sh` with mock agents
- Intercept agent command structure inside mock `herdr`
- Verify agent receives correct conversation UUID flag
- Invalid/missing UUID recovery path test
- Workspace resume CLI args verification | | **Stop Session** | - Session name verification check
- Purge verification confirmations logic
- Command derivation format validation
- Timeout calculation helper tests
- Reason logging serializer test | - Safe folder path validation (shutil protection)
- Database status field mutation serialization
- Isolation folder cleanup check
- Lock file release checks on stop
- Concurrency handling of stop mutations | - Execute `stop_session.sh` with graceful key delivery (`/exit`)
- Fallback to forcible termination (`herdr kill-session`) check
- Fallback to PID termination (`kill -9`) verify
- Purge files verification on disk (`--purge-conversation`)
- CLI flag verification with yes/no confirmation | | **Status Query** | - JSON converter unit tests
- Diff formatter text generators
- Output alignment tests
- Table grid column math verify
- CLI status argument parse tests | - Status read locks verification
- Parsing of drift status classifications
- Concurrency read protection test
- Registry YAML-to-JSON structural translation
- Verification of database read access checks | - Running `status.sh` with `--json`
- Verify console output match formatting rules
- Verify exit status codes on different states
- Integration test with `reconcile.sh` read-only diff emission
- Verify status command doesn't trigger side effects | | **Monitor/Reconcile** | - Drift state classification unit tests
- Signature verification checks
- Subscription topic parsing tests
- MQTT message structure validator
- HMAC validation logic tests | - Concurrency lock checks (`.mam/monitor.lock`)
- Verify YAML and SQLite database reconciliation logic
- DB validation on drift updates
- HMAC signature signature verification
- SQLite journal mode fallback check (WAL vs DELETE) | - Execute `reconcile.sh` in single-pass mode (`--once`)
- MQTT subscription execution with mock messages
- Verify auto-termination of orphaned herdr sessions
- Verify auto-registration of untracked herdr sessions
- Lock contention handling testing | ## Test Architecture The E2E testing framework is built using **pytest** and relies on two main pillars to ensure hermetic and reproducible test runs: 1. **Environment Sandboxing**: All tests run inside a temporary, isolated directory structure provided by the pytest `tmp_path` fixture. The workspace environment is sandboxed by: - Creating a temporary `.mam/` directory. - Using the `monkeypatch` fixture to override `AGENT_SESSIONS_YAML` pointing to the sandboxed path. - Overriding relevant environment variables (like `HOME`, `WORKSPACE_ROOT`, etc.) to prevent tests from modifying the developer's system state. 2. **Mock Binaries Interception**: To prevent tests from interacting with external systems or relying on running daemons: - A mock `herdr` script is dynamically generated and placed in a temporary bin folder, which is prepended to the system `PATH`. This mock binary reads/writes to a JSON file (`mock_herdr_state.json`) which acts as the control pane for tests to assert that `herdr` was called with correct arguments and return mocked outputs (session list, capture-pane output, exit codes). - Mock agent binaries (`claude`, `agy`, `hermes`, `cline`) are also generated and prepended to `PATH`. They emulate successful login verification commands (e.g. `claude auth status`) and mock conversation UUID generation on disk. ## Real-World Application Scenarios (Tier 4) We define five key E2E scenarios representing end-to-end user workflows: 1. **Standard Agent Session Lifecycle**: Spawning a new worker agent session via `create_session.sh`, verifying it is registered correctly in the YAML database, checking its status via `status.sh`, and then gracefully stopping it via `stop_session.sh`. 2. **Session Disconnect and Resume**: Creating a session, simulating a network disconnect/agent pane termination (updating herdr state), calling `resume_session.sh` to restore it using the workspace-scoped UUID, and asserting that the session returns to the active state in both herdr and the registry. 3. **Drift Detection and Auto-Reconciliation**: Artificially introducing drift (e.g. terminating a herdr session manually from the backend while keeping it registered in the YAML registry, or starting a herdr session outside the scripts), running `reconcile.sh --once`, and verifying that orphaned sessions are terminated and registry state is updated. 4. **Parallel Session Operations with flock Locking**: Simulating concurrent creation/stop script invocations to verify that SQLite flock transactions block lost update races, and that the registry data remains consistent. 5. **Multi-Agent Orchestrator Review Loop**: Running the orchestrator loop (`run_loop.sh`) where a worker agent and a reviewer agent are spawned, reviewer verdicts (`PASS` and `NOT PASS`) are processed, loops are iterated, and planner escalation is triggered on failure. ## Coverage Thresholds To ensure the test suite is comprehensive, we define the following coverage thresholds: - **Tier 1 (Unit Tests)**: Minimum >=5 unit tests per feature (total >=25 unit tests). - **Tier 2 (Component Tests)**: Minimum >=5 component tests per feature (total >=25 component tests). - **Tier 3 (Integration Tests)**: Pairwise combination testing covering CLI options and environment overrides for all features. - **Tier 4 (E2E Scenarios)**: At least 5 full real-world scenario tests implemented and passing.