Install any skill in seconds. Free to start, no credit card required.
Get Started Free →帮助用户撰写高质量的文献综述类论文。提供从选题、文献检索、评估筛选、结构规划到最终写作的全流程指导。适用于需要撰写独立文献综述论文或学术论文中文献综述部分的用户。
.claude/skills/brycewang-stanford-literature-review/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 100% | 5 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -23% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 3% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 38% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 47% | 0% |
This is an original Open Science Skills workflow for experimental and computational social science. It remixes high-level ideas from Cheng-I Wu's Academic Research Skills for Claude Code (CC BY-NC 4.0), especially evidence mapping, source verification, and mode separation between narrative literature review and formal systematic review. It is not a full ARS pipeline and should not copy ARS prose.
Decide what the user needs:
Default to a narrative/evidence-map review unless the user explicitly asks for a systematic review, meta-analysis, or PRISMA-compliant output.
Before summarizing papers, specify:
If the user only gives a broad topic, first produce a short scoping memo with 2-4 possible review boundaries rather than writing a generic review.
Use the user's supplied sources first. Then identify obvious missing source classes:
Run citation-check when the source list is large, messy, DOI-heavy, or likely to contain stale working papers.
For each important source, record:
Do not produce chronological "Author A says X, Author B says Y" prose unless chronology is theoretically important.
Organize sources into 3-6 clusters. Prefer conceptual or mechanism clusters over method-only clusters:
For each cluster, state what is settled, what is contested, and what would change the interpretation.
Write a gap verdict:
When the gap is weak, propose a better contribution frame rather than only criticizing it.
narrative-building after the evidence map exists to turn the review into the "Why-to-If-Then" funnel.hypothesis-building when the review implies falsifiable expectations and estimands.pre-registration-writing when the review supports confirmatory hypotheses.methods-reporting when reviewing how prior studies report designs, sample flow, and transparency.journal-review when auditing someone else's manuscript for novelty and placement.Produce a Literature Review Evidence Map:
# Literature Review Evidence Map
Review question:
Scope and exclusions:
Search/source base:
Gap verdict: Holds / Partly holds / Does not hold / Cannot assess
## Closest Prior Work
| Source | What it actually establishes | Boundary | Relation to user's claim |
## Evidence Clusters
### Cluster 1: <name>
Settled:
Contested:
Missing:
Key sources:
## Contribution Diagnosis
Claimed gap:
Verdict:
Better contribution frame:
## Literature Review Architecture
1. <section purpose>
2. <section purpose>
3. <section purpose>
## Sentences the Review Must Earn
- <sentence-level claim that needs source support>
## Sources Needing Verification
| Source | Why |citation-check was invoked or recommended when source integrity was uncertain.narrative-building.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-10 | pass→pass | 16,265 | 18,355 | +13% | 1 | 1 | 0% | 3,136 | 4,803 | +53% | 0 | 0 | — |
case-09 | pass→pass | 19,791 | 25,127 | +27% | 1 | 1 | 0% | 4,377 | 6,559 | +50% | 0 | 0 | — |
case-01 | fail→pass | 39,466 | 30,778 | -22% | 1 | 1 | 0% | 6,304 | 7,237 | +15% | 0 | 0 | — |
case-02 | fail→pass | 33,737 | 24,151 | -28% | 1 | 1 | 0% | 6,257 | 4,836 | -23% | 0 | 0 | — |
case-03 | fail→pass | 31,625 | 26,627 | -16% | 1 | 1 | 0% | 6,257 | 6,466 | +3% | 0 | 0 | — |
case-04 | pass→pass | 15,211 | 13,164 | -13% | 1 | 1 | 0% | 3,064 | 3,175 | +4% | 0 | 0 | — |
case-05 | fail→pass | 17,842 | 19,205 | +8% | 1 | 1 | 0% | 3,103 | 4,288 | +38% | 0 | 0 | — |
case-06 | fail→fail | 20,294 | 18,908 | -7% | 1 | 1 | 0% | 4,465 | 5,039 | +13% | 0 | 0 | — |
case-07 | fail→pass | 15,873 | 16,536 | +4% | 1 | 1 | 0% | 3,129 | 4,585 | +47% | 0 | 0 | — |
case-08 | pass→pass | 23,119 | 16,082 | -30% | 1 | 1 | 0% | 3,397 | 4,320 | +27% | 0 | 0 | — |
case-11 | fail→pass | 16,139 | 28,921 | +79% | 1 | 1 | 0% | 2,914 | 5,508 | +89% | 0 | 0 | — |
case-12 | fail→pass | 15,556 | 14,817 | -5% | 1 | 1 | 0% | 2,857 | 4,180 | +46% | 0 | 0 | — |
case-13 | fail→pass | 17,709 | 20,155 | +14% | 1 | 1 | 0% | 3,439 | 4,865 | +41% | 0 | 0 | — |
case-14 | pass→pass | 16,748 | 24,331 | +45% | 1 | 1 | 0% | 2,875 | 4,816 | +68% | 0 | 0 | — |
case-15 | fail→fail | 15,177 | 24,407 | +61% | 1 | 1 | 0% | 2,708 | 5,219 | +93% | 0 | 0 | — |
case-16 | fail→pass | 26,751 | 14,662 | -45% | 1 | 1 | 0% | 3,707 | 3,654 | -1% | 0 | 0 | — |
case-17 | pass→pass | 14,125 | 22,066 | +56% | 1 | 1 | 0% | 2,566 | 5,157 | +101% | 0 | 0 | — |
case-18 | fail→pass | 23,052 | 20,845 | -10% | 1 | 1 | 0% | 4,086 | 5,052 | +24% | 0 | 0 | — |
case-19 | fail→pass | 18,666 | 20,426 | +9% | 1 | 1 | 0% | 3,337 | 4,893 | +47% | 0 | 0 | — |
case-20 | fail→pass | 26,176 | 22,847 | -13% | 1 | 1 | 0% | 3,535 | 4,655 | +32% | 0 | 0 | — |
case-21 | fail→fail | 24,328 | 32,486 | +34% | 1 | 1 | 0% | 6,214 | 7,483 | +20% | 0 | 0 | — |
case-22 | fail→fail | 31,190 | 36,925 | +18% | 1 | 1 | 0% | 5,577 | 6,592 | +18% | 0 | 0 | — |
case-23 | fail→fail | 6,160 | 25,383 | +312% | 1 | 1 | 0% | 1,038 | 5,603 | +440% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +52 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.