Files
multi-agent-mux/.agents/reports/reviewer-cline-01/report-825cb977.md
T
Godopu 4bbd03bf2d fix(lib): handle agent_not_ready startup state and reject claude fullscreen upsell modal
- Accept agent_not_ready from herdr agent start to allow startup dialog handling without premature rollback, while preserving fail-closed behavior on dead process timeouts.
- Match Claude fullscreen renderer upsell modal via 'Yes, try it' and dismiss with Escape to avoid dropping permission flags or deadlocking on idle /tui tips.
- Add behavioral test suite in test_b19_headless_reconcile_fixes.py and cross-agent review reports.
2026-08-27 10:17:53 +09:00

7.2 KiB
Raw Blame History

Cross-Code Review Report — Job 825cb977

Reviewer: reviewer-cline-01 (Cline) Changeset: .agents/skills/lib.sh (+12/4), tests/test_b19_headless_reconcile_fixes.py (+106), FIX.md (new, 22 lines). Scope: lint, behavior, and loss/orphan cross-review of the two lib.sh fixes + accumulated diff.

1. Are the targeted problems real? (discard gate)

The brief's stated "goal" text describes the original FIX.md intent (allow timed out waiting for agent startup + add generic fullscreen tokens). The actual diff does the corrected opposite on point 1 and a safer variant on point 2. Both addressed problems are real — this is not a discard candidate.

  • Fix 1 — agent-start detection. Real problem: the prior code treated only agent_started as success, so herdr's documented agent_not_ready ("process up, blocked on a dialog") status caused rollback of a legitimately-starting agent that merely needed dialog handling. The fix promotes agent_not_ready to success (→ wait_for_tui_ready) and classifies fatal CLI errors first. It also excludes timed out waiting for agent startup from success — correct, because that string is ambiguous (herdr returns it for a dead /bin/false too), so promoting it would misclassify a dead process and waste the 30s readiness window.
  • Fix 2 — fullscreen renderer upsell modal. Real problem: Claude's fullscreen upsell modal (Yes, try it) is a blocking dialog. The fix adds the modal-unique token Yes, try it to _MAM_DIALOG_TOKENS and dismisses with Escape (reject). This is the safe choice: Enter would accept Yes, try it and restart the session without --dangerously-skip-permissions (permission-flag drop). The idle /tui fullscreen tip (Try the new fullscreen renderer … · /tui fullscreen, with a prompt) is intentionally not matched — it is non-blocking, and a generic fullscreen renderer token would false-match that ready idle screen and deadlock wait_for_tui_ready.

2. Lint

  • bash -n .agents/skills/lib.sh → OK.
  • .venv/bin/python -m py_compile tests/test_b19_headless_reconcile_fixes.py → OK.
  • Shell quoting/regex consistent with surrounding code: fatal-error grep -qiE (case-insensitive ERE) first; success grep -qE "agent_started|agent_not_ready" (literal alternation, no unescaped metachars); new Yes, try it token is a literal with no ERE specials — safe inside grep -Eq/grep -q.
  • No shellcheck-style issues introduced (no unquoted expansions, no word-splitting hazards in the added lines).

3. Behavior

  • Fix 1 (lib.sh L546-556): fatal errors (^usage:/^error:/etc.) break with success=0 → downstream if [ "$success" -ne 1 ] (L562) → exit 1 (fail-fast). agent_started|agent_not_readysuccess=1; break → proceeds to wait_for_tui_ready. Timeout-only output → no match → retries (3 backoffs ≈3.5s) → exit 1 (fast dead-process failure instead of a 30s wait). The success init/check chain is intact.
  • Fix 2 (lib.sh L62, L1859-1862): Yes, try it added to _MAM_DIALOG_TOKENS (so _pane_dialog_open detects the modal — also correctly gates send_keys_safe against prompting under a modal) and to handle_startup_dialogs (sends Escape). The branch is placed before Yes, proceed and the readiness-token branch — correct ordering (modal must be dismissed before ready detection). After Escape the loop re-captures and returns 0 once the banner appears; bounded by timeout (default 20s). The idle tip contains no Yes, try it_pane_dialog_open returns false → wait_for_tui_ready detects the banner (no deadlock).
  • Tests: the 4 new tests are genuine behavior tests (stub _pane_capture/_sks_herdr/sleep, source the real lib.sh, exercise real _pane_dialog_open/handle_startup_dialogs/wait_for_tui_ready). _LIB_SH uses Path(__file__).resolve() (CWD-independent). One source-string guard (test_agent_start_success_tokens_exclude_startup_timeout) asserts token membership + error-before-success ordering.

4. Loss / Orphan analysis

  • lib.sh: the removed standalone if grep -q "agent_started"; then success=1; break; fi is fully superseded by the combined agent_started|agent_not_ready check — no orphaned variable or branch. success=0 init and the downstream success-ne-1 guard remain consistent. The new Yes, try it→Escape branch is self-contained; no existing branch was orphaned.
  • Tests: from pathlib import Path is used by _LIB_SH; both _FULLSCREEN_TIP/_FULLSCREEN_MODAL fixtures are used; _run_lib_helpers is used by 3 behavior tests. No unused imports or dead helpers introduced.
  • No lost functionality: agent_not_ready is a superset-preserving addition (still proceeds to wait_for_tui_ready); the timeout exclusion is an intentional, justified narrowing (ambiguous token), not a loss of needed behavior. FIX.md is an accurate working note (untracked, expected to ship with the fix).

5. Test results

  • Cited 4 suites (test_b19_headless_reconcile_fixes.py, test_herdr_shim_contract.py, test_a4_adapter_contract.py, test_b8_send_keys_verification.py) → 29 passed.
  • Broader sweep pytest tests/ -q: ~378 tests passed with 0 failures (full unit + component + tier1/2 + tier3 integration all green). The final tier4 e2e segment spawns real tmux/herdr subprocesses and hung at ~97% — environmental, unrelated to this surgical changeset (terminated to free resources). Zero failure lines in the output.
  • Regression-guard effectiveness (mutation-tested in the prior adjudication pass on this same diff, re-confirmed here by inspection): Escape→Enter on the modal makes test_fullscreen_modal_is_rejected_not_accepted FAIL; re-adding fullscreen renderer to _MAM_DIALOG_TOKENS makes test_fullscreen_tip_is_not_a_blocking_dialog FAIL. Guards are non-vacuous.

6. Edge cases examined

  • E-1: Branch order in handle_startup_dialogsYes, try it precedes Yes, proceed and the readiness branch. The two dialogs are distinct (no token overlap); order is safe and correct (dismiss modal before ready).
  • E-2: Other consumers of _MAM_DIALOG_TOKENSsend_keys_safe gating via _pane_dialog_open also treats the modal as a dialog (blocks prompting under a modal). Consistent and desirable.
  • E-3: Yes, try it false-positive risk — specific affirmative phrase unique to the upsell modal; the tip fixture (contains Try the new fullscreen renderer but not Yes, try it) returns DIALOG_CLOSED. Low risk; acceptable.
  • E-4: Fatal-error regex ^error: (case-insensitive) ordered first — if herdr ever emitted both an error line and a status, fatal wins (fail-safe). herdr success outputs are status lines, not error:. No conflict.
  • E-5: agent_not_ready→success then wait_for_tui_ready — if the process is up but never shows a banner (unhandled dialog), the readiness loop is bounded (30s) → abort. No zombie.
  • E-6: Escape on the modal re-captures next iteration; if the banner appears → return 0; if the modal re-appeared (unlikely) it would Escape again, bounded by the 20s timeout. Safe.

7. Verdict

Both targeted problems are real and correctly fixed. The changeset is surgical, lint-clean, behavior-tested with non-vacuous guards, and introduces no orphans or lost functionality. No design-level rework is required.

[VERDICT: PASS]