Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Deep literature review — PubMed, EuropePMC, bioRxiv preprints, citation networks, evidence synthesis. Disambiguates queries, runs collision-aware searches, grades evidence T1-T4, and produces structured reports. Use for systematic literature review, meta-analysis evidence collection, and detailed answer-with-citations workflows.
.claude/skills/mims-harvard-tooluniverse-literature-deep-research/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-12 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 94% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 75% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 276% | 0% |
Systematic literature research: disambiguate, search with collision-aware queries, grade evidence, produce structured reports.
KEY PRINCIPLES: (1) Disambiguate first (2) Right-size deliverable (3) Grade every claim T1-T4 (4) All sections mandatory even if "limited evidence" (5) Source attribution for every claim (6) English-first queries, respond in user's language (7) Report = deliverable, not search log
Search PubMed/EuropePMC FIRST before reasoning. A published paper beats memory.
Factoid search strategy:
EuropePMC_search_articles(query="term1 term2 term3", limit=5)When analysis requires computation (statistics, data processing, scoring, enrichment), write and run Python code via Bash. Don't describe what you would do — execute it and report actual results. Use ToolUniverse tools to retrieve data, then Python (pandas, scipy, statsmodels, matplotlib) to analyze it.
Phase 0: Clarify + Mode Select → Phase 1: Disambiguate + Profile → Phase 2: Literature Search → Phase 3: Report| Mode | When | Deliverable | |------|------|-------------| | Factoid | Single concrete question | 1-page fact-check report + bibliography | | Mini-review | Narrow topic | 1-3 page narrative | | Full Deep-Research | Comprehensive overview | 15-section report + bibliography |
markdown# [TOPIC]: Fact-check Report ## Question / ## Answer (with evidence rating) / ## Source(s) / ## Verification Notes / ## Limitations
| Pattern | Domain | Action | |---------|--------|--------| | Gene/protein symbol | Biological target | Full bio disambiguation | | Drug name | Drug | Drug disambiguation (1.5) | | Disease name | Disease | Disease disambiguation (1.6) | | CS/ML topic | General academic | Skip bio tools, literature-only | | Cross-domain | Interdisciplinary | Resolve each entity in its domain |
tooluniverse-target-researchtooluniverse-drug-researchtooluniverse-disease-researchUse this skill for literature synthesis. Use specialized skills for entity profiling. For max depth, run both.
UniProt_search → UniProt_get_entry_by_accession → UniProt_id_mapping
ensembl_lookup_gene → MyGene_get_gene_annotationCheck first 20 results. If >20% off-topic, build negative filter: NOT [collision1] NOT [collision2]. Gene family: "ADAR" NOT "ADAR2" NOT "ADARB1". Cross-domain: add context terms.
InterPro_get_protein_domains, UniProt_get_ptm_processing_by_accession, HPA_get_subcellular_location,
GTEx_get_median_gene_expression, GO_get_annotations_for_gene, Reactome_map_uniprot_to_pathways,
STRING_get_protein_interactions, intact_get_interactions, OpenTargets_get_target_tractability_by_ensemblIDGPCR targets: delegate to tooluniverse-target-research.
Identity: OpenTargets_get_drug_chembId_by_generic_name, ChEMBL_get_drug, PubChem_get_CID_by_compound_name, drugbank_get_drug_basic_info_by_drug_name_or_id Targets: ChEMBL_get_drug_mechanisms, OpenTargets_get_associated_targets_by_drug_chemblId, DGIdb_get_drug_gene_interactions Safety: OpenTargets_get_drug_adverse_events_by_chemblId, OpenTargets_get_drug_indications_by_chemblId, search_clinical_trials
OpenTargets disease search → EFO/MONDO IDs
DisGeNET_get_disease_genes, DisGeNET_search_disease
CTD_get_disease_chemicalsResolve both entities, then cross-reference via CTD_get_chemical_gene_interactions, CTD_get_chemical_diseases, OpenTargets drug-target/drug-disease tools. Intersect shared targets/pathways.
Non-bio: skip bio tools, use ArXiv/DBLP/OSF. Cross-domain: resolve bio entities with 1.1-1.3, search CS/general in parallel, merge and cross-reference.
Methodology stays internal. Report shows findings, not process.
Step 1: Seeds (15-30 core papers): domain-specific title searches with date/sort filters. Step 2: Citation expansion: PubMed_get_cited_by, EuropePMC_get_citations/references, PubMed_get_related, SemanticScholar_get_recommendations, OpenCitations_get_citations. If the opt-in Noodle MCP is connected (noodle_*, needs NOODLE_MCP_URL), its bounded citation/semantic graph traversal is another angle on the same PubMed corpus -- a discovery signal, not evidence of causality or validity, same caveat as the others. Step 3: Collision-filtered broader queries: "[TERM]" AND ([context]) NOT [collision]
Run the core multi-field set on every review (catches what any single index misses), then add the domain rows that match the subject. Don't fire every source blindly — 6–10 well-chosen indexes beat 20 noisy ones.
ALWAYS run (core, all disciplines): PubMed_search_articles, EuropePMC_search_articles, openalex_search_works (query param search/query) or openalex_literature_search (query param search_keywords) — pick one and match its param; mixing them silently returns off-topic results — and SemanticScholar_search_papers
Then add by domain:
| Domain | Add these | Notes | |--------|-----------|-------| | Biomedical / clinical | PMC_search_papers (full text), PubTator3_LiteratureSearch (entity & relations: queries), PubMed_Guidelines_Search (clinical guidelines) | PubTator normalizes gene/drug/disease entities | | Biology (ecology/evolution/plant) | EuropePMC as PRIMARY + OpenAlex | PubMed returns 0–1 for non-clinical biology | | CS / ML / AI | ArXiv_search_papers, DBLP_search_publications | arXiv + CS bibliography | | Physics / HEP / astro | InspireHEP_search_papers | 1.6M+ particle/astro records | | Broad / hard-to-find / OA | Crossref_search_works, CORE_search_papers, DOAJ_search_articles, Fatcat_search_scholar, Consensus_search_papers | DOI registry + OA aggregators + Internet Archive Scholar; Consensus (220M+ papers) adds an AI takeaway + study-design metadata per paper -- useful for fast triage, not a substitute for reading the source | | Regional / EU-funded | OpenAIRE_search_publications, HAL_search_archive | EU open science + French national archive | | Datasets / software / outputs | Figshare_search_articles, Zenodo_search_records | Citable DOIs for data & code | | Preprints (latest) | EuropePMC_search_articles(source='PPR'), OSF_search_preprints, BioRxiv_get_preprint/MedRxiv_get_preprint (DOI lookup) | bioRxiv/medRxiv/PsyArXiv etc. |
Multi-source: advanced_literature_search_agent (12+ DBs; needs Azure key -- fallback: query the core set individually). Citation impact: iCite_search_publications (RCR/APT), iCite_get_publications (by PMID), scite_get_tallies (support/contradict). PubMed-only; for CS use SemanticScholar.
A domain-specific index returning 0 (e.g. ArXiv on a pure-clinical topic) is normal — only worry if the whole core set is empty.
Full-text: see FULLTEXT_STRATEGY.md for three-tier strategy.
CRITICAL: PubMed returns 0 for ~30% of valid queries. Always retry with EuropePMC when PubMed returns empty. This is not optional.
Retry once -> fallback tool. Key fallbacks: PubMed_get_cited_by -> EuropePMC_get_citations -> OpenCitations. OA: Unpaywall if configured, else Europe PMC/PMC/OpenAlex flags.
Last resort when every structured index above is empty (a brand-new preprint, a dataset page, a project site with no DOI): the opt-in exa_* tools (general neural web search, no key needed for casual use) can still find it, but it's general internet retrieval, not a scientific database -- verify anything it surfaces against a real source before citing, don't grade it T1-T4 as if it were literature.
| Tier | Label | Bio Example | CS/ML Example | |------|-------|-------------|---------------| | T1 | Mechanistic | CRISPR KO + rescue, RCT | Formal proof, controlled ablation | | T2 | Functional | siRNA knockdown phenotype | Benchmark with baselines | | T3 | Association | GWAS, screen hit | Observational, case study | | T4 | Mention | Review article | Survey, workshop abstract |
Inline: Target X regulates Y [T1: PMID:12345678]. Per theme: summarize evidence distribution.
Triaging a large candidate set before reading in full: Consensus_search_papers returns study type and sample size per paper, a fast first pass for provisional tiering -- confirm against the actual paper before citing, its metadata is a starting point, not the grade itself.
| File | Mode | |------|------| | [topic]_report.md | Full | | [topic]_factcheck_report.md | Factoid | | [topic]_bibliography.json + .csv | All |
Progressive update: create report with all section headers immediately. Fill after each phase. Write Executive Summary LAST.
Use 15-section template from REPORT_TEMPLATE.md. Domain adaptations: bio (architecture/expression/GO/disease), drug (properties/MOA/PK/safety), disease (epi/patho/genes/treatments), general (history/theories/evidence/applications).
Brief progress updates only: "Resolving identifiers...", "Building paper set...", "Grading evidence..." Do NOT expose: raw tool outputs, dedup counts, search round details.
TOOL_NAMES_REFERENCE.md -- 130+ tools with parametersREPORT_TEMPLATE.md -- template, domain adaptations, bibliography, completeness checklistFULLTEXT_STRATEGY.md -- three-tier full-text verificationWORKFLOW.md -- compact cheat-sheetEXAMPLES.md -- worked examples| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 16,226 | 32,071 | +98% | 1 | 1 | 0% | 2,987 | 9,127 | +206% | 0 | 0 | — |
case-02 | fail→fail | 34,398 | 10,708 | -69% | 1 | 1 | 0% | 6,202 | 3,195 | -48% | 0 | 0 | — |
case-03 | fail→fail | 36,662 | 10,059 | -73% | 1 | 1 | 0% | 6,191 | 3,288 | -47% | 0 | 0 | — |
case-04 | pass→fail | 10,884 | 6,571 | -40% | 1 | 1 | 0% | 2,169 | 3,023 | +39% | 0 | 0 | — |
case-05 | pass→fail | 10,561 | 125,537 | +1089% | 1 | 1 | 0% | 2,086 | 2,971 | +42% | 0 | 0 | — |
case-06 | pass→fail | 18,001 | 9,377 | -48% | 1 | 1 | 0% | 3,389 | 3,167 | -7% | 0 | 0 | — |
case-12 | fail→pass | 17,223 | 6,422 | -63% | 1 | 1 | 0% | 3,072 | 3,729 | +21% | 0 | 0 | — |
case-07 | fail→fail | 10,191 | 8,681 | -15% | 1 | 1 | 0% | 1,884 | 3,143 | +67% | 0 | 0 | — |
case-08 | fail→pass | 13,954 | 14,169 | +2% | 1 | 1 | 0% | 2,255 | 4,385 | +94% | 0 | 0 | — |
case-09 | pass→fail | 14,809 | 6,711 | -55% | 1 | 1 | 0% | 2,587 | 3,051 | +18% | 0 | 0 | — |
case-10 | pass→pass | 15,431 | 11,266 | -27% | 1 | 1 | 0% | 2,586 | 4,611 | +78% | 0 | 0 | — |
case-11 | pass→fail | 9,623 | 7,689 | -20% | 1 | 1 | 0% | 1,434 | 3,077 | +115% | 0 | 0 | — |
case-13 | pass→fail | 17,103 | 36,463 | +113% | 1 | 1 | 0% | 2,931 | 3,305 | +13% | 0 | 0 | — |
case-14 | fail→fail | 12,845 | 51,028 | +297% | 1 | 1 | 0% | 2,188 | 2,962 | +35% | 0 | 0 | — |
case-15 | fail→pass | 16,350 | 3,895 | -76% | 1 | 1 | 0% | 2,549 | 3,304 | +30% | 0 | 0 | — |
case-16 | pass→pass | 7,367 | 4,360 | -41% | 1 | 1 | 0% | 1,183 | 3,240 | +174% | 0 | 0 | — |
case-17 | pass→pass | 8,941 | 14,362 | +61% | 1 | 1 | 0% | 1,580 | 4,144 | +162% | 0 | 0 | — |
case-18 | pass→fail | 17,106 | 14,152 | -17% | 1 | 1 | 0% | 3,103 | 4,139 | +33% | 0 | 0 | — |
case-19 | fail→pass | 15,635 | 12,537 | -20% | 1 | 1 | 0% | 2,598 | 4,542 | +75% | 0 | 0 | — |
case-20 | fail→pass | 5,468 | 2,659 | -51% | 1 | 1 | 0% | 800 | 3,011 | +276% | 0 | 0 | — |
case-21 | fail→pass | 12,472 | 5,039 | -60% | 1 | 1 | 0% | 2,016 | 3,389 | +68% | 0 | 0 | — |
case-22 | pass→fail | 14,265 | 7,687 | -46% | 1 | 1 | 0% | 2,367 | 3,109 | +31% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 9 counted toward the lift figure. The other 13 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -9 percentage points is the difference between those two pass rates over the 9 comparable cases. 10 cases got worse with the skill loaded, and they are included in that figure.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.