Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Conducting systematic, scoping, and narrative literature reviews. Covers PRISMA/PRISMA-ScR protocols, search strategy (Boolean, MeSH), database selection (PubMed, Scopus, Web of Science, Embase), screening, data extraction, evidence synthesis (narrative, meta-analysis, thematic), and reporting. Use when planning or executing a formal literature review.
.claude/skills/jaechang-hits-literature-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 167% | 0% |
| case-16 | ✓→✗ | ▼ Worse | 146% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 171% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 202% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 191% | 0% |
A literature review systematically identifies, appraises, and synthesizes published evidence on a defined research question. The method ranges from informal narrative reviews to highly structured systematic reviews with meta-analysis. Choosing the correct review type, building a reproducible search strategy, and applying transparent inclusion/exclusion criteria are the foundational decisions that determine whether a review can be trusted and published in a high-impact journal. This guide covers the full workflow from question formulation to synthesis and reporting.
| Review Type | Definition | When to Use | Time Required | |-------------|-----------|-------------|--------------| | Narrative review | Selective, expert-curated synthesis; no protocol; no PRISMA | Introducing a topic; describing mechanistic background | Days to weeks | | Scoping review | Comprehensive mapping of evidence landscape; PRISMA-ScR; no quality appraisal | Understand what evidence exists before committing to systematic review | Weeks to months | | Systematic review | Exhaustive search; predefined protocol (PROSPERO); quality appraisal; PRISMA | Answer a specific clinical/scientific question with highest rigor | Months to years | | Meta-analysis | Systematic review + quantitative pooling of effect estimates | Quantify pooled effect size and heterogeneity across studies | Months to years | | Umbrella review | Systematic review of existing systematic reviews | Synthesize evidence from multiple reviews on one topic | Months | | Rapid review | Streamlined systematic review with time-limited methods | Time-sensitive policy or clinical decisions | Weeks |
Peer reviewer expectations: Journals in the biomedical domain expect systematic and scoping reviews to follow PRISMA or PRISMA-ScR reporting standards and to be pre-registered in PROSPERO (systematic reviews only). Narrative reviews are typically invited by editors rather than submitted unsolicited.
Systematic reviews require a precisely defined research question. The PICO framework structures the question into searchable, operationalizable components:
| Component | Meaning | Example (for a clinical question) | |-----------|---------|----------------------------------| | P | Population | Adults with type 2 diabetes aged ≥ 40 | | I | Intervention | SGLT2 inhibitors (empagliflozin, dapagliflozin) | | C | Comparison | Placebo or standard care | | O | Outcome | Cardiovascular mortality, HbA1c, eGFR decline | | S | Study design (PICOS) | Randomized controlled trials only |
For basic science questions, adapt to PECO (Population, Exposure, Comparator, Outcome) or a custom framework. A well-formed PICO directly maps to search terms for each database.
Different study designs provide different levels of certainty about causal effects. Standard hierarchy for intervention questions (highest to lowest):
Systematic reviews and meta-analyses of RCTs (highest certainty)
↓
Individual randomized controlled trials (RCTs)
↓
Non-randomized controlled trials / quasi-experiments
↓
Prospective cohort studies
↓
Retrospective cohort / case-control studies
↓
Cross-sectional studies
↓
Case series and case reports
↓
Expert opinion / narrative review / editorials (lowest certainty)For diagnostic accuracy, prognosis, and etiology questions, the hierarchy differs. The GRADE framework (Grading of Recommendations Assessment, Development and Evaluation) formalizes evidence quality across four domains: risk of bias, inconsistency, indirectness, and imprecision.
No single database covers all literature. Major databases and their coverage:
| Database | Coverage | Strength | Access | |----------|----------|----------|--------| | PubMed / MEDLINE | >35M biomedical records; 1946+ | Free; MeSH controlled vocabulary; high precision | Free | | Embase | >34M records; European + drug focus; 1947+ | Best for pharmacology and European journals | Subscription | | Web of Science | ~90M records across science and humanities | Citation analysis; interdisciplinary | Subscription | | Scopus | ~90M records; broad | Largest abstract database; good non-English | Subscription | | CINAHL | Nursing and allied health | Best for nursing/PT/OT research | Subscription | | PsycINFO | Psychology and behavioral science | Deep coverage of behavioral literature | Subscription | | Cochrane CENTRAL | Controlled trials only; curated | Highest precision for RCTs | Free/subscription | | ClinicalTrials.gov | US-registered clinical trials | Includes unpublished/ongoing trials | Free | | Grey literature | Reports, theses, guidelines | Reduces publication bias | Various |
Systematic reviews should search at minimum: PubMed + Embase + one domain-specific database + Cochrane CENTRAL (for clinical topics). Searching only PubMed biases toward US/English publications and misses up to 30% of relevant trials.
What is your literature review goal?
│
├── "Understand the topic background for my paper's Introduction"
│ └── → Narrative review: select key papers; no protocol needed
│
├── "Map what evidence exists before designing a study"
│ └── → Scoping review (PRISMA-ScR): comprehensive but no quality appraisal
│
├── "Answer a specific clinical or scientific question rigorously"
│ ├── Quantitative data poolable across studies?
│ │ ├── Yes → Systematic review + meta-analysis (PRISMA + PROSPERO)
│ │ └── No → Systematic review with narrative synthesis only
│ └── Time-constrained (policy deadline)?
│ └── → Rapid review (document scope limitations)
│
└── "Synthesize existing systematic reviews"
└── → Umbrella review| Review type | Pre-registration required? | Quality appraisal? | Reporting standard | Minimum databases | |-------------|---------------------------|-------------------|--------------------|------------------| | Narrative | No | No | None formal | Author's choice | | Scoping | No (recommended) | No | PRISMA-ScR | ≥2 major databases | | Systematic | Yes (PROSPERO) | Yes (RoB 2, ROBINS-I, etc.) | PRISMA 2020 | ≥3 databases | | Meta-analysis | Yes (PROSPERO) | Yes | PRISMA 2020 | ≥3 databases | | Umbrella | Recommended | Yes (AMSTAR-2) | PRISMA | ≥2 databases |
citation-management — reference managers for collecting and organizing search results before and during reviewstatistical-analysis — statistical methods for meta-analysis (pooled effects, heterogeneity, forest plots)scientific-critical-thinking — evaluating individual study quality and interpreting effect sizes in the context of a review| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 13,139 | 27,449 | +109% | 1 | 1 | 0% | 2,093 | 5,676 | +171% | 0 | 0 | — |
case-02 | fail→fail | 12,619 | 10,541 | -16% | 1 | 1 | 0% | 1,980 | 5,046 | +155% | 0 | 0 | — |
case-03 | pass→pass | 9,788 | 11,164 | +14% | 1 | 1 | 0% | 1,737 | 5,248 | +202% | 0 | 0 | — |
case-04 | pass→pass | 11,714 | 13,680 | +17% | 1 | 1 | 0% | 1,948 | 5,663 | +191% | 0 | 0 | — |
case-05 | pass→pass | 12,834 | 9,929 | -23% | 1 | 1 | 0% | 2,254 | 5,081 | +125% | 0 | 0 | — |
case-06 | fail→pass | 10,168 | 10,140 | -0% | 1 | 1 | 0% | 1,815 | 4,854 | +167% | 0 | 0 | — |
case-07 | pass→pass | 10,220 | 9,174 | -10% | 1 | 1 | 0% | 1,723 | 4,946 | +187% | 0 | 0 | — |
case-12 | pass→pass | 10,893 | 9,292 | -15% | 1 | 1 | 0% | 1,832 | 4,875 | +166% | 0 | 0 | — |
case-08 | pass→pass | 11,111 | 10,539 | -5% | 1 | 1 | 0% | 1,941 | 5,299 | +173% | 0 | 0 | — |
case-09 | pass→pass | 10,502 | 10,606 | +1% | 1 | 1 | 0% | 1,717 | 5,188 | +202% | 0 | 0 | — |
case-10 | pass→pass | 8,720 | 7,240 | -17% | 1 | 1 | 0% | 1,645 | 4,698 | +186% | 0 | 0 | — |
case-11 | pass→pass | 15,488 | 13,607 | -12% | 1 | 1 | 0% | 2,628 | 5,687 | +116% | 0 | 0 | — |
case-13 | pass→pass | 11,047 | 7,690 | -30% | 1 | 1 | 0% | 1,718 | 4,577 | +166% | 0 | 0 | — |
case-14 | pass→pass | 12,383 | 14,473 | +17% | 1 | 1 | 0% | 2,108 | 5,839 | +177% | 0 | 0 | — |
case-15 | pass→pass | 12,051 | 11,741 | -3% | 1 | 1 | 0% | 2,070 | 5,251 | +154% | 0 | 0 | — |
case-16 | pass→fail | 13,742 | 13,983 | +2% | 1 | 1 | 0% | 2,271 | 5,583 | +146% | 0 | 0 | — |
case-17 | pass→pass | 5,267 | 2,961 | -44% | 1 | 1 | 0% | 1,003 | 3,887 | +288% | 0 | 0 | — |
case-18 | pass→pass | 10,787 | 5,571 | -48% | 1 | 1 | 0% | 1,861 | 4,405 | +137% | 0 | 0 | — |
case-19 | pass→pass | 12,716 | 9,369 | -26% | 1 | 1 | 0% | 2,144 | 5,005 | +133% | 0 | 0 | — |
case-20 | pass→pass | 12,026 | 13,192 | +10% | 1 | 1 | 0% | 2,609 | 6,153 | +136% | 0 | 0 | — |
case-21 | pass→pass | 9,686 | 18,553 | +92% | 1 | 1 | 0% | 1,835 | 5,681 | +210% | 0 | 0 | — |
case-22 | pass→pass | 3,675 | 3,226 | -12% | 1 | 1 | 0% | 717 | 3,971 | +454% | 0 | 0 | — |
case-23 | pass→pass | 9,261 | 10,869 | +17% | 1 | 1 | 0% | 1,734 | 5,426 | +213% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of 0 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.