feat(agent): deprecate and completely remove cline agent support
- Delete adapters/cline.py and unregister from registry.py - Remove cline branches from lib.sh and all 8 skill scripts (create, resume, stop, status, reconcile, update_yaml_resumed, resolve_session_id, orc_onboard) - Narrow own-key mapping dictionaries across lib_py core modules to 4 supported agents - Delete cline-exclusive tests and retarget shared fixtures to grok/hermes/claude - Update skills documentation and installation guides (439 passed, 0 failures) - Archive cline deprecation consensus and review reports
This commit is contained in:
@@ -0,0 +1,18 @@
|
||||
# Report: Job e30b9201 — Refined Cline Removal Plan (Rev.2) per `creator-agy-01` Challenge
|
||||
|
||||
**Durable output (updated in place)**: [.agents/reports/planner-reviewer-claude-01/plan-264c3b5d.md](../../../.agents/reports/planner-reviewer-claude-01/plan-264c3b5d.md)
|
||||
|
||||
## Summary
|
||||
|
||||
`creator-agy-01` challenged Rev.1's test-retargeting strategy (§5) and doc-scope decision (§6), identifying one core blind spot and two supporting gaps. **All accepted — no `[REBUT:]` filed**, after independently re-verifying each claim:
|
||||
|
||||
1. **Core blind spot**: Rev.1 proposed retargeting the 2-tier readiness tests by pre-exporting `MAM_STRONG_READY_TOKENS`/`MAM_WEAK_READY_TOKENS` env vars directly. I re-checked `lib.sh` myself and confirmed this would skip `wait_for_tui_ready`'s `python -m lib_py.agents facts <agent>` bridge call entirely (it's gated behind an `[ -z "$strong_tokens" ]` check) — since cline was the only real adapter with non-degenerate 2-tier values, this would leave the bridge's handling of genuinely-differentiated strong/weak tokens completely untested once cline is gone. **Fixed**: retarget to a test-local mock adapter instead, keeping the real facts bridge in the loop.
|
||||
2. **`test_c8` gap**: Rev.1 only swapped the modal *text*, missing that the test's session name (`my-project-worker-cline-01`) drives dynamic agent resolution → `facts cline` — a synthetic modal string wouldn't fix that the session name itself would fail to resolve once cline's adapter is deleted. **Fixed**: retarget the whole test to `claude`, using claude's own real `modal_tokens` value (`Try the new fullscreen renderer?`, verified present in `claude.py`), preserving the complete E2E path with an agent that isn't being removed.
|
||||
3. Added a new adapter-contract assertion (all 4 remaining agents' default 2-tier fields round-trip correctly through the facts bridge) to close the coverage gap explicitly.
|
||||
4. **`docs/NEW_AGENT_INTEGRATION_GUIDE.md`**: I'd placed this in "do not touch," but verified myself it's a live copy-paste template with `cline` hardcoded into 5 sample sites (adapter dict, `lib.sh` kind mapping, binary tuple, diagram, test assertion) — moved into the update-scope in §6.
|
||||
|
||||
Full Rev.2 text with the corrected test-retargeting specifications is in the durable plan linked above.
|
||||
|
||||
**Process note**: still pending a fresh `[VERDICT: PASS]` before implementation begins, per `MULTI_AGENT_RULES.md` §3.1.
|
||||
|
||||
[VERDICT: N/A — planning artifact, review pending]
|
||||
@@ -0,0 +1,25 @@
|
||||
# Report: Job adad47d1 — Refined Consensus (Rev.2) per `creator-agy-01` Challenge
|
||||
|
||||
**Durable output (updated in place)**: [.agents/reports/cline_deprecation_opinion.md](../../../.agents/reports/cline_deprecation_opinion.md)
|
||||
|
||||
## Summary
|
||||
|
||||
`creator-agy-01` challenged Rev.1's §3.1 risk characterization of cline's unattended-execution gap, arguing it understated the real danger: `cline history --json` (the pre-spawn gate) only validates that local session storage exists, not credential/API validity — so it can't protect against a runtime auth failure (401/expired token/429 quota) mid-task, which could plausibly force cline's TUI into its interactive setup modal with no clean-exit fallback like the other four agents have. A second point noted the 40-session usage fact I'd used to support RETAIN also amplifies `discover()`'s session-hijacking exposure.
|
||||
|
||||
**Verification before accepting**: I ran `cline history --help` myself (confirmed: pure local listing, no auth semantics) and traced exactly where MAM checks for the setup modal in `lib.sh` (confirmed: only at spawn-time and at the moment of injecting a new prompt — never continuously during an autonomous work stretch). Both confirm the architectural gap is real. I was also careful to flag what remains *inferred* rather than *observed*: neither `creator-agy-01` nor I have actually triggered a live credential failure to watch cline's real behavior — I adopted the escalated risk model as the working precautionary assumption while saying so explicitly, rather than overstating certainty either side doesn't have.
|
||||
|
||||
**Disposition**: Accepted, with the honesty caveat above — no `[REBUT:]` filed.
|
||||
|
||||
## What changed in Rev.2
|
||||
|
||||
- §3.1 rewritten: risk escalated from "narrow spawn-time race" to "unmonitored mid-task modal-fallback window," with the verification trail (what's confirmed vs. inferred) laid out explicitly.
|
||||
- §3.2 narrowed to reaffirm only the tool-call-approval finding that still stands (`--auto-approve` default true) — the credential-failure gap is a separate, more severe category, not covered by that flag.
|
||||
- New §3.4: precision-scoped the session-hijacking concern (cline's `discover()` does filter by workspace `cwd`, so exposure is concurrent-same-workspace activity, not all 40 sessions indiscriminately) while accepting the core point.
|
||||
- §4 revised: my RETAIN vote stands but its terms changed — now explicitly conditional on the new §6.2 safety gate, not an unconditional RETAIN as Rev.1 implied.
|
||||
- §6 reprioritized: the reconcile.sh fix is now urgent; a new mandatory safety-gate requirement (exclude cline from long-running autonomous roles until upstream ships a real headless/non-interactive flag) is added for the RETAIN path; the empirical-verification ask is upgraded from "nice to have" to a prerequisite before trusting cline with new unattended work.
|
||||
|
||||
Full Rev.2 text is in the durable report linked above.
|
||||
|
||||
**Process note**: Still no unanimous verdict (2 RETAIN-conditional vs. 1 REMOVE) — this challenge sharpened the RETAIN side's conditions rather than flipping any vote. The REMOVE-vs-RETAIN decision still belongs to the General Manager per `MULTI_AGENT_RULES.md` §3.
|
||||
|
||||
[VERDICT: N/A — consensus opinion artifact, decision belongs to the General Manager]
|
||||
@@ -0,0 +1,40 @@
|
||||
# 🔍 Cross-Code Review — Cline Deprecation Consensus Opinion, Rev.3 (Job 073e27d4)
|
||||
|
||||
- **Reviewer**: `planner-reviewer-claude-01`
|
||||
- **Target document**: `.agents/reports/cline_deprecation_opinion.md` (Rev.3 — response to `reviewer-creator-grok-01` (job `0a056794`) and `reviewer-hermes-01` (job `bc68ed65`) reviews)
|
||||
- **Prior review context**: I also reviewed Rev.2 of this same document in job `b9a72dce`, flagging 3 findings (imprecise `§6.2`/`§4.2` cross-references, an overclaimed "proved ~0.5–1.5 days" statement, and a stale §1 summary line). Rev.3 does not appear to be a response to that review — its changelog (F1/F2/F3) addresses different findings from grok's and hermes's separate reviews. This review checks both: whether Rev.3's own new claims are accurate, and whether my earlier findings were carried forward.
|
||||
- **Method**: read the full current file (109 lines, confirmed via `wc -l`; `git status --short` shows only this file as untracked, confirming the "no skill/framework code modified" constraint holds), and independently re-verified Rev.3's three new corrections (F1/F2/F3) against the live `cline --help` output and `reconcile.sh` source rather than accepting the changelog's claims at face value.
|
||||
|
||||
---
|
||||
|
||||
## 1. Constraint Compliance
|
||||
|
||||
`git status --short` → only `?? .agents/reports/cline_deprecation_opinion.md`. No skill/framework code touched. Diff header claims `+109` lines; live file is 109 lines — consistent.
|
||||
|
||||
## 2. Verification of Rev.3's Own New Claims (F1/F2/F3)
|
||||
|
||||
I did not take the changelog's self-description at face value — I re-derived each claim independently:
|
||||
|
||||
- **F1 (flag inventory)**: Ran `cline --help` myself. Confirmed line 25 of its output: `-k, --key <api-key> API key override for this run`. The report's bounded framing — this flag injects a key at startup but cannot refresh a credential mid-task or suppress the interactive modal fallback on a runtime provider failure, and no `--headless`/`--non-interactive` flag exists — is accurate; I found nothing in `cline --help` contradicting that scope-limiting claim. **Verified correct.**
|
||||
- **F2 (drift-C modernization status, corrected line numbers)**: I grepped `reconcile.sh` for `sibling_claimed` and drift-C block headers. Initially my grep for `"drift C ("` missed claude's block because its header uses a different format (`# === drift C: claude ...` — colon, not a parenthesized agent name, unlike agy/hermes/cline's `# === drift C (agy): ...` style). On closer inspection, claude's block **is** at line 637 exactly as claimed, and it indeed calls `verify_session_uuid(cwd, 'claude', uuid, s, mode="discover")` with the raw row `s` — no `sibling_claimed` exclusion, matching cline's block at line 785. agy (line 692) and hermes (line 742) both build `s_eval['_sibling_claimed_uuids']`. **Verified correct** — this is a genuine improvement over Rev.1/Rev.2, which had incorrectly implied claude's block was already modernized (grouping it with agy/hermes).
|
||||
- **F3 (consensus attribution)**: Cross-checked against hermes's original report (`.mam/jobs/57f33eff/hermes-reports/report-final.md`), which does state the drift-C fix as an explicit numbered condition of its RETAIN verdict, and grok's report, which lists the drift-C block within cline's maintenance-cost inventory (to be deleted under REMOVE) rather than as a standalone precondition. **Verified correct.**
|
||||
|
||||
All three of Rev.3's own corrections are accurate and represent genuine, verified improvements over Rev.2.
|
||||
|
||||
## 3. Findings Carried Forward — Unaddressed from My Rev.2 Review (job `b9a72dce`)
|
||||
|
||||
Rev.3's changelog responds to grok's and hermes's reviews, but none of the three issues I flagged in my own separate Rev.2 review were incorporated. Re-verified as still present in the live Rev.3 text:
|
||||
|
||||
- **Still present** (line 14, 85): `§6.2` cited as if it were a subsection heading. §6 (line 102) is still a flat `## 6.` heading followed by a plain numbered list (`1.`, `2.`) — no `### 6.1`/`### 6.2` headings exist anywhere in the document. Same issue as before, unfixed.
|
||||
- **Still present** (line 49): "the grok integration already **proved** the reverse operation (adding an agent) costs ~0.5–1.5 days" — unchanged. As I found in the prior review, `grok.py` was added in a single squashed commit (`ad8201d`, 2026-08-26), which cannot establish actual wall-clock effort; the day-count traces to `new_agent_types_roadmap.md`'s a priori estimate for different hypothetical candidates, not a measured fact about grok. "Proved" still overstates this.
|
||||
- **Still present** (line 34): "**2 of 3 lean RETAIN** (both conditional on the same follow-up fix)" — unchanged. As of Rev.2, my own RETAIN vote already carried an additional condition (the §6 item 2 safety gate) that hermes's original report never agreed to, so the two RETAIN votes are not conditioned on literally the same thing. This has been true since Rev.2 and remains uncorrected in Rev.3.
|
||||
|
||||
## 4. Minor Observation (not a defect)
|
||||
|
||||
The §0 changelog's Rev.2 entry was compressed from Rev.1/Rev.2's original 4-row table (which included per-point verification methodology, e.g. "I ran `cline history --help` myself") into shorter prose bullets during the Rev.3 restructuring. Some audit-trail granularity was lost from the changelog summary specifically, though the underlying detail still lives in the body sections (§3.1, etc.) it refers to. Not a correctness issue, just a slight reduction in the changelog's own self-sufficiency as a summary.
|
||||
|
||||
## 5. Verdict
|
||||
|
||||
Rev.3's own corrections (F1/F2/F3) are all independently verified accurate and are genuine improvements — in particular, F2 correctly identifies that claude's drift-C block is just as un-modernized as cline's, which earlier revisions had gotten wrong. However, three previously-identified, still-valid findings from my prior review of this same document were not carried forward into this revision. None of these — old or new — are severe enough to undermine the document's core methodology or conclusions; they remain small, mechanical precision fixes. Passing, with the expectation that a future revision finally closes out all outstanding findings from both review passes together rather than only the most recent one.
|
||||
|
||||
[VERDICT: PASS]
|
||||
@@ -0,0 +1,48 @@
|
||||
# 🔍 Cross-Code Review — Complete Cline Removal Implementation (Job 20d45d12)
|
||||
|
||||
- **Reviewer**: `planner-reviewer-claude-01`
|
||||
- **Target diff**: implementation of `plan-264c3b5d.md` Rev.2 — 30 files (adapter deletion, registry, 4 `lib_py` modules, `lib.sh`, 9 skill scripts, 6 test files, 3 docs) plus 1 out-of-scope test-flakiness fix.
|
||||
- **Method**: read every changed file's live post-diff state directly (not diff text alone), independently verified the two highest-risk items from my own Rev.2 plan (the `reconcile.sh` drift-C block boundary and the tiered-readiness/modal test retargeting), syntax-checked all 10 modified shell scripts, grepped the entire diff for any surviving `cline` reference, and ran the full test suite myself.
|
||||
|
||||
---
|
||||
|
||||
## 1. Fidelity to Rev.2 Plan — Verified, Not Assumed
|
||||
|
||||
I did not trust the implementation's own claim of compliance — I re-checked the specific corrections `creator-agy-01`'s challenge required in Rev.2 against the live diff:
|
||||
|
||||
- **Tiered-readiness tests (Rev.2's core correction)**: `test_c3_strong_token_and_hint_token_readiness_succeeds` now mocks `_delegate_py_bin`/`python -m lib_py.agents facts` to return a synthetic `mocktiered` agent with genuinely distinct `MAM_STRONG_READY_TOKENS='MockApp'` / `MAM_WEAK_READY_TOKENS='Use arrow keys'`, keeping the real facts-bridge call path exercised rather than bypassing it with raw env-var pre-injection — exactly what Rev.2 required. The other C4–C7 tests that were already using pre-set env vars (not resolving through the bridge in the original cline-based version either) were correctly left as simple session-name swaps, since they were never testing the bridge to begin with.
|
||||
- **`test_c8b` (modal test)**: retargeted fully to `claude`, using `'Try the new fullscreen renderer?'` as the injected screen text and `MAM_MODAL_TOKENS='Try the new fullscreen renderer\?'`, with session name `my-project-worker-claude-01` — this is claude's actual, verified `modal_tokens` value, exactly matching Rev.2's requirement to preserve the full session-name-resolution → `facts claude` → dialog-block path, not a synthetic placeholder.
|
||||
- **New facts-bridge round-trip assertion**: `test_a4_adapter_contract.py::test_adapter_required_properties` gained `assert adapter.strong_ready_tokens == adapter.ready_tokens` / `assert adapter.weak_ready_tokens == ''` for all 4 remaining agents, and `test_facts_bridge_eval_contract` now additionally asserts `STRONG=`/`WEAK=` come through the real bash `eval` of the bridge's output — this is actually a **stronger** implementation than what I asked for (I only required the property-level check; this round-trips through the real subprocess + bash eval too).
|
||||
- **`docs/NEW_AGENT_INTEGRATION_GUIDE.md`**: all 5 sites I flagged in Rev.2 (diagram, `_ADAPTERS` sample, `lib.sh` kind-mapping sample, binary-tuple sample, test-assertion sample) were updated — the architecture diagram box-drawing was even correctly realigned (`┬` connector fixed) after swapping `ClineAgentAdapter` for `GrokAgentAdapter` in that slot, not just text-deleted.
|
||||
|
||||
## 2. Independent Verification of the Highest-Risk Edit
|
||||
|
||||
I flagged the `reconcile.sh` cline drift-C block deletion as the highest-risk single edit in my own plan. Checked the live file directly: the block is cleanly gone, the preceding `hermes` drift-C block and the following `result = {...}` return statement are both intact and correctly adjacent with no orphaned fragments. Extracted and `ast.parse()`'d the actual `RECON_SRC` heredoc (lines 320–794, not the other heredoc earlier in the file, which I made sure to distinguish) — valid Python. `bash -n` on the whole file — valid.
|
||||
|
||||
## 3. Completeness Check
|
||||
|
||||
`git diff | grep -n "^+.*[Cc]line"` (every added line, across the entire diff) returns **zero matches** — no newly-written line anywhere in this diff still references cline. Cross-checked a full-repo `cline` grep against `git status`: every remaining match is either inside `.agents/reports/**` (untouched, correct) or inside changelog-style docs (`VERSIONS.md`, `IMPROVEMENTS.md`) describing past releases in the past tense (correctly left alone, consistent with my plan's "spot-check, don't blanket-edit" guidance).
|
||||
|
||||
## 4. Findings
|
||||
|
||||
### 4.1 Minor: `MULTI_AGENT_RULES.md`/`.ko.md` line 21 slightly stale (Low, not blocking)
|
||||
|
||||
`"Newly spawned agents (e.g., antigravity, claude, cline, hermes) act as Team Leaders..."` — an illustrative `e.g.` list, not a hard enumeration, but it does still name cline as a live example post-removal. Low severity since the sentence's substance is about the *role concept*, not a supported-agent contract, and this file wasn't in either of our removal plans' scope. Worth a follow-up touch-up, not blocking.
|
||||
|
||||
### 4.2 Out-of-scope change present in the diff (informational, not a defect)
|
||||
|
||||
`tests/test_o2_race_free_lock.py` was modified — replacing a fixed `time.sleep(0.3)` in `acquire_bg()` with an active poll-until-marker-file-written loop (up to 2s, with early exit if the background process dies). This has nothing to do with cline removal; it's a flaky-test timing fix, most likely surfaced while chasing "100% pass, zero regressions" during implementation. I reviewed the change itself: it's strictly safer than what it replaces (removes a fixed-sleep race assumption, fails faster on a dead process) and doesn't touch cline-adjacent code. Flagging for transparency/scope-discipline reasons, not as a defect — I would not block on this alone.
|
||||
|
||||
## 5. Full Test Suite
|
||||
|
||||
```
|
||||
.venv/bin/python -m pytest tests/ -q
|
||||
→ 439 passed in 655.22s (0:10:55), exit code 0
|
||||
```
|
||||
Ran to completion myself (not the diff's own claim). **Zero failures, zero regressions.**
|
||||
|
||||
## 6. Verdict
|
||||
|
||||
Every site from my own Rev.2 plan was implemented faithfully and, in two places (the facts-bridge round-trip assertion, the architecture-diagram realignment), more thoroughly than the plan strictly required. No orphaned `cline` references anywhere in the diff. The highest-risk edit (`reconcile.sh`'s block deletion) is clean and syntactically valid. One low-severity doc staleness and one out-of-scope-but-safe test fix are noted, neither blocking.
|
||||
|
||||
[VERDICT: PASS]
|
||||
Reference in New Issue
Block a user