Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when a mostly complete ML conference paper needs self-review, pre-submission QA, camera-ready checking, section-by-section critique, citation-risk inspection, or rebuttal/review-response drafting. Skip this for initial drafting and use `paperreview` only when the user explicitly wants external submission.
.claude/skills/cnfjlhj-paper-review-pipeline/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-17 | ✗→✓ | ▲ Improved | 69% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 111% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 98% | 0% |
| case-21 | ✗→✓ | ▲ Improved | -53% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 116% | 0% |
Run a two-view paper review for ML conference submissions:
1) Section-by-section review (Abstract → Intro → Method → Experiments → …) with concrete edits. 2) Prioritized issue list with P0/P1/P2 severity, grouped by category, including recommended fixes and verification notes.
This skill also supports rebuttal / review response: parse reviewer comments, classify, choose a strategy, and draft a professional point-by-point response.
This skill is a consolidation layer. It must not omit any distinctive workflow, constraints, or output formats from the legacy skills it replaces.
Use:
references/parity-matrix.md as the feature-parity contract and regression scenarios.references/modules/ for full imported workflows and checklists.This skill supports two modes:
targeted — run only the most relevant tracks based on the user request and inputs, but always produce the Final Synthesis.full-parallel — run all tracks as independent outputs (acceptable redundancy), then produce the Final Synthesis.Trigger full-parallel when the user says: “全量”, “并行”, “run all tracks”, “run every skill”, “逐个 skill 测试”, “full pipeline”.
Protocol + required report template:
references/full-parallel-protocol.mdreferences/report-template.mdpaperreview)If the user provides a near-final or final PDF, complete the local review first and then explicitly ask whether they also want an external second opinion via paperreview.
paperreview as a follow-up step; otherwise finish with the local pipeline result only.1) No hallucinated citations.
[CITATION NEEDED] / placeholder and tell the user explicitly.2) Do not change technical meaning.
3) Preserve LaTeX semantics when editing source.
\\cite{}, \\ref{}, \\label{}, math environments, figures/tables, or bibliography hooks.4) Respect blind review constraints (if applicable).
For each section, output:
Use the checklists in:
references/section-review-checklist.mdFormat each issue like:
Use the taxonomy in:
references/p0-p2-taxonomy.mdUse these modules to preserve legacy feature parity:
references/modules/paper-self-review.mdreferences/modules/review-response.md and references/rebuttal-workflow.mdreferences/modules/academic-paper-helper.mdreferences/modules/citation-validator.md and references/citation-integrity.mdreferences/modules/writing-anti-ai.mdreferences/modules/latex-rhythm-refiner.mdreferences/modules/claude-scholar-ml-paper-writing.md1) Identify the one-sentence contribution and confirm it with the user. 2) Extract the top 3–7 claims the paper relies on. 3) For each claim, note the current support:
4) Flag immediate P0 risks (typical: missing baselines, unclear experimental protocol, unverified citations, paper “about X” but experiments test Y).
Review sections in order (and check alignment between them): 1) Abstract 2) Introduction (motivation → gap → contribution bullets) 3) Related Work (positioning + not a bibliography dump) 4) Method (reproducible description + design justification) 5) Experiments / Results (fair baselines + full setup + statistical reporting) 6) Analysis / Ablations (claim-driven, not exploratory noise) 7) Limitations / Broader Impact / Ethics (venue-dependent) 8) Conclusion (tight restatement + constraints + future work without overclaim)
Turn findings into an actionable issue list:
If the user wants an execution plan, produce:
When mode is full-parallel, do not collapse everything into one voice. Instead:
1) Run all tracks (A–G) as independent “mini-reviewers” and keep each track output visible. 2) Then produce the Final Synthesis:
Use:
references/full-parallel-protocol.mdreferences/report-template.mdWhen reviews arrive, do: 1) Parse and classify each comment: Major / Minor / Clarification / Missing baseline / Missing experiment / Writing / Citation / Misunderstanding. 2) Choose a strategy per item: Accept + change, Clarify, Defend, Add experiment, Defer (explain constraints). 3) Draft point-by-point responses with:
Use:
references/rebuttal-workflow.mdBefore submission, ensure:
If deep citation audit is requested, follow:
references/citation-integrity.md| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 12,646 | 38,411 | +204% | 1 | 1 | 0% | 1,939 | 8,421 | +334% | 0 | 0 | — |
case-02 | fail→fail | 9,888 | 13,261 | +34% | 1 | 1 | 0% | 1,541 | 4,229 | +174% | 0 | 0 | — |
case-03 | fail→fail | 5,998 | 12,106 | +102% | 1 | 1 | 0% | 806 | 4,172 | +418% | 0 | 0 | — |
case-04 | fail→fail | 6,699 | 5,953 | -11% | 1 | 1 | 0% | 917 | 3,063 | +234% | 0 | 0 | — |
case-05 | fail→fail | 7,019 | 7,128 | +2% | 1 | 1 | 0% | 1,028 | 3,203 | +212% | 0 | 0 | — |
case-06 | pass→pass | 14,728 | 13,994 | -5% | 1 | 1 | 0% | 2,193 | 4,299 | +96% | 0 | 0 | — |
case-07 | pass→pass | 28,188 | 14,917 | -47% | 1 | 1 | 0% | 2,101 | 4,318 | +106% | 0 | 0 | — |
case-08 | pass→pass | 10,247 | 12,700 | +24% | 1 | 1 | 0% | 1,754 | 4,327 | +147% | 0 | 0 | — |
case-09 | pass→pass | 9,509 | 9,343 | -2% | 1 | 1 | 0% | 1,547 | 3,551 | +130% | 0 | 0 | — |
case-10 | fail→fail | 7,330 | 3,689 | -50% | 1 | 1 | 0% | 1,037 | 2,682 | +159% | 0 | 0 | — |
case-11 | fail→fail | 8,813 | 4,920 | -44% | 1 | 1 | 0% | 1,325 | 2,874 | +117% | 0 | 0 | — |
case-12 | fail→fail | 7,961 | 11,274 | +42% | 1 | 1 | 0% | 1,167 | 3,810 | +226% | 0 | 0 | — |
case-17 | fail→pass | 14,036 | 11,948 | -15% | 1 | 1 | 0% | 2,322 | 3,918 | +69% | 0 | 0 | — |
case-13 | fail→pass | 11,814 | 9,558 | -19% | 1 | 1 | 0% | 1,739 | 3,664 | +111% | 0 | 0 | — |
case-14 | pass→pass | 14,058 | 11,018 | -22% | 1 | 1 | 0% | 2,115 | 3,740 | +77% | 0 | 0 | — |
case-15 | pass→pass | 13,114 | 10,217 | -22% | 1 | 1 | 0% | 1,973 | 3,779 | +92% | 0 | 0 | — |
case-16 | fail→fail | 6,946 | 6,017 | -13% | 1 | 1 | 0% | 1,110 | 2,980 | +168% | 0 | 0 | — |
case-18 | pass→pass | 5,856 | 7,720 | +32% | 1 | 1 | 0% | 1,057 | 3,300 | +212% | 0 | 0 | — |
case-19 | fail→fail | 3,172 | 6,548 | +106% | 1 | 1 | 0% | 402 | 3,079 | +666% | 0 | 0 | — |
case-20 | fail→pass | 9,435 | 5,381 | -43% | 1 | 1 | 0% | 1,465 | 2,905 | +98% | 0 | 0 | — |
case-21 | fail→pass | 28,859 | 4,355 | -85% | 1 | 1 | 0% | 5,901 | 2,793 | -53% | 0 | 0 | — |
case-22 | fail→pass | 17,977 | 9,511 | -47% | 1 | 1 | 0% | 1,614 | 3,488 | +116% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +23 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.