Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Mechanize Pattern 15 — the seven-pass adversarial review protocol for academic manuscripts. Spawns 7 forked subagents in parallel (abstract, intro, methods, results, robustness, prose, citations), then synthesizes a prioritized revision checklist. Use for submission-ready or R&R-stage papers where single-pass review isn't enough.
.claude/skills/pedrohcgs-seven-pass-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 128% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 59% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 191% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 94% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 65% | 0% |
Runs seven independent reviewers, each focused on a single lens, then synthesizes their findings into one prioritized revision plan — the fan-out → reduce → judge runtime from orchestrator-protocol.md, applied with seven lenses.
Why seven passes? A single-agent review blends lenses and softens each one. Seven forked agents each approach the paper with full context budget for their own lens, then a synthesizer resolves conflicts and de-duplicates.
> When to pick this over /review-paper: This skill costs roughly 7× more tokens than /review-paper (default) and ~2× more than /review-paper --adversarial. Use it when the paper is submission-ready or at R&R stage and you need maximum lens coverage. For early drafts or iterative work, /review-paper is the right tool. For journal-simulation pressure test, use /review-paper --peer <journal> instead.
$0 — manuscript path (.tex, .qmd, .md, or .pdf). Required.Each lens runs as a forked subagent (context: fork) so the main conversation stays clean.
| # | Lens | Focus | Agent type | |---|---|---|---| | 1 | Abstract audit | Does the abstract state the question, method, result, and contribution? Does it match the paper? | general-purpose | | 2 | Intro structure | Does the intro follow Cochrane / Varian framework? Literature placement? Contribution clarity? | general-purpose | | 3 | Methods / identification | Are assumptions stated? Is identification credible? Are alternatives addressed? | domain-reviewer | | 4 | Results + tables | Do tables read standalone? Is magnitude + significance discussed? Units consistent? | general-purpose | | 5 | Robustness | Are obvious threats pre-empted? Is the robustness section convincing or theatrical? | general-purpose | | 6 | Prose quality | Sentence-level clarity, hedging, passive voice, paragraph cohesion | proofreader | | 7 | Citation audit | Invokes /validate-bib --semantic; checks cite-claim direction for top-10 works | general-purpose |
.pdf → extract text first (pdftotext -layout).quality_reports/seven_pass_[stem]/.In a single message, spawn 7 Agent tool calls (one per lens). Each subagent gets:
quality_reports/seven_pass_[stem]/lens_[N]_[lens-name].md.finding-schema.json, written to quality_reports/seven_pass_[stem]/lens_[N]_[lens-name].json beside the prose report. Severities: blocker | major | minor | nit. Every finding computes its id with python3 scripts/validate-findings.py --id FILE LINE LOCUS and carries rule, evidence, and a failing_case. Phase 2 validates each array first (python3 scripts/validate-findings.py <file> — exit 0 required; a lens whose report does not validate has not reviewed), then reduces over the typed findings — it does not re-read the prose. Because ids are lens-independent, the same defect found by two lenses dedups to one finding automatically.This is the fan-out primitive from orchestrator-protocol.md; Agent subagents are the portable mechanism (the agents that fill lenses 3/6 are in agent-fleet.md).
Lens prompt rubrics are embedded inline below — one summary paragraph per lens. Each forked subagent receives its lens's rubric plus the manuscript path.
Lens prompt summaries:
/validate-bib --semantic. For top-10 cited works, does the in-text claim match the cited paper's actual finding direction? Are contemporary / competing works cited?Wait for all 7 lens reports. Reduce, don't re-review: stack the seven scorecards and apply the gate predicate from orchestration-schemas.md §3 — the Executive verdict is a function of the typed findings, not a fresh eighth opinion. Then run the post-judge hallucination gate (§4): any CRITICAL the synthesis introduces that no lens raised must be re-verified in a fresh claim-verifier fork, or dropped to [JUDGE-HALLUCINATED] and the verdict recomputed. A synthesis may freely downgrade or de-duplicate lens findings; it may not invent a new blocker.
Then produce:
quality_reports/seven_pass_[stem]/_SYNTHESIS.md
markdown# Seven-Pass Review: [Manuscript] **Date:** YYYY-MM-DD **Path:** [manuscript] ## Executive verdict **Overall state:** [SUBMIT / REVISE-MINOR / REVISE-MAJOR / REJECT-AND-RESTART] ## Cross-lens CRITICAL issues | # | Lens(es) | Issue | Recommendation | |---|---|---|---| ## MAJOR issues (second-round) | # | Lens(es) | Issue | |---|---|---| ## MINOR polish [bulleted] ## Per-lens scorecard | Lens | Critical | Major | Minor | Score/10 | |---|---|---|---|---| | 1. Abstract | | | | | | 2. Intro | | | | | | 3. Methods | | | | | | 4. Results | | | | | | 5. Robustness | | | | | | 6. Prose | | | | | | 7. Citations | | | | | | **Overall** | | | | | ## Revision plan (in recommended order) 1. [Highest-leverage fix — usually a lens with 2+ CRITICALs] 2. … 7. [Lowest-leverage polish] ## Contradictions between lenses [If two lenses disagree, surface here. E.g., Lens 2 says "expand contribution" but Lens 6 says "trim intro".]
After synthesis, print:
Seven-pass review complete.
Subagents: 7 (parallel) + 1 synthesizer.
Approx token usage: ~80–120k (vs ~15k for single-pass /review-paper).
Runtime: ~3–5 min wall-clock.
For cheaper alternatives:
- Single-pass: /review-paper
- Iterative: /review-paper --adversarial/review-paper single-pass first).This skill's reviewers emit findings under the machine-checked contract in finding-schema.json. Reports are JSON arrays.
Smoke-test the harness before spending review effort — a run that fans out reviewers and then cannot write a valid report has wasted the whole pass:
bashecho '[]' | python3 scripts/validate-findings.py
Then, before presenting any summary:
bashpython3 scripts/validate-findings.py <report>.json # exit 0 required
What the contract forces, and why:
rule — the documented rule or standard violated. A finding citing no rule is anopinion, and opinions do not gate a commit.
failing_case — a concrete configuration under which the claim breaks, or the exactmissing hypothesis. "This could be clearer" does not validate.
id = sha1("<file>:<line>:<locus>") — deterministic, so dedup across rounds isexact and the two-strikes rule is checkable rather than eyeballed.
mechanical — true only for fixes that cannot change a result (typo, cross-reference,formatting, label). Never for an estimand, assumption, specification, inference procedure, sample definition, or reporting language: those return to the researcher.
Apply the per-lens evidence burdens and the "does NOT count" filters in orchestration-schemas.md §7 before verification, so known false alarms never reach the judge. The verifier pass is refute-biased: only verdict: "confirmed" findings ship; anything it cannot ground is dropped, not downgraded to a warning.
.claude/skills/review-paper/SKILL.md — the single-pass and --adversarial modes (cheaper, faster)..claude/skills/validate-bib/SKILL.md — invoked by Lens 7..claude/skills/audit-reproducibility/SKILL.md — complementary; numeric-claims side of the audit.CRITICAL at the top of the synthesis should block submission until resolved._SYNTHESIS.md, skip unchanged lenses if requested via --incremental (future)./review-paper --adversarial's job.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 35,007 | 66,961 | +91% | 1 | 1 | 0% | 375 | 3,261 | +770% | 0 | 0 | — |
case-02 | fail→fail | 48,023 | 26,791 | -44% | 1 | 1 | 0% | 8,296 | 3,386 | -59% | 0 | 0 | — |
case-03 | fail→fail | 34,905 | 7,080 | -80% | 1 | 1 | 0% | 369 | 3,407 | +823% | 0 | 0 | — |
case-04 | fail→pass | 11,180 | 4,471 | -60% | 1 | 1 | 0% | 1,515 | 3,456 | +128% | 0 | 0 | — |
case-05 | fail→fail | 7,225 | 23,383 | +224% | 1 | 1 | 0% | 570 | 6,063 | +964% | 0 | 0 | — |
case-06 | fail→fail | 8,784 | 4,774 | -46% | 1 | 1 | 0% | 1,036 | 2,965 | +186% | 0 | 0 | — |
case-07 | fail→pass | 15,106 | 8,689 | -42% | 1 | 1 | 0% | 2,544 | 4,051 | +59% | 0 | 0 | — |
case-08 | fail→pass | 6,435 | 2,018 | -69% | 1 | 1 | 0% | 1,044 | 3,034 | +191% | 0 | 0 | — |
case-09 | fail→fail | 12,639 | 4,348 | -66% | 1 | 1 | 0% | 2,143 | 3,478 | +62% | 0 | 0 | — |
case-10 | pass→pass | 4,869 | 3,653 | -25% | 1 | 1 | 0% | 749 | 3,332 | +345% | 0 | 0 | — |
case-11 | pass→pass | 4,380 | 2,381 | -46% | 1 | 1 | 0% | 640 | 3,051 | +377% | 0 | 0 | — |
case-12 | fail→pass | 10,922 | 2,916 | -73% | 1 | 1 | 0% | 1,628 | 3,151 | +94% | 0 | 0 | — |
case-13 | fail→pass | 14,283 | 4,903 | -66% | 1 | 1 | 0% | 2,144 | 3,545 | +65% | 0 | 0 | — |
case-14 | pass→pass | 8,229 | 3,235 | -61% | 1 | 1 | 0% | 1,400 | 3,272 | +134% | 0 | 0 | — |
case-15 | fail→pass | 9,552 | 1,986 | -79% | 1 | 1 | 0% | 1,536 | 2,982 | +94% | 0 | 0 | — |
case-16 | fail→pass | 10,311 | 3,978 | -61% | 1 | 1 | 0% | 1,881 | 3,411 | +81% | 0 | 0 | — |
case-17 | fail→pass | 13,956 | 1,830 | -87% | 1 | 1 | 0% | 2,527 | 2,986 | +18% | 0 | 0 | — |
case-18 | fail→pass | 18,233 | 1,855 | -90% | 1 | 1 | 0% | 3,079 | 3,002 | -3% | 0 | 0 | — |
case-19 | fail→pass | 10,432 | 6,608 | -37% | 1 | 1 | 0% | 1,864 | 3,963 | +113% | 0 | 0 | — |
case-20 | fail→pass | 8,205 | 4,210 | -49% | 1 | 1 | 0% | 1,308 | 3,471 | +165% | 0 | 0 | — |
case-21 | fail→pass | 10,254 | 2,672 | -74% | 1 | 1 | 0% | 1,676 | 3,178 | +90% | 0 | 0 | — |
case-22 | fail→pass | 7,369 | 2,326 | -68% | 1 | 1 | 0% | 1,121 | 3,134 | +180% | 0 | 0 | — |
case-23 | pass→pass | 7,463 | 2,908 | -61% | 1 | 1 | 0% | 1,228 | 3,256 | +165% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 19 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +57 percentage points is the difference between those two pass rates over the 19 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/13/2026 | +64% |
Other measured skills in the registry, with their headline benchmark lift.