Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Multi-agent system for biomedical literature review and synthesis
.claude/skills/brycewang-stanford-med-researcher-guide/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -39% | 0% |
| case-18 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-01 | ✗→✓ | ▲ Improved | -20% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 13% | 0% |
Med-Researcher is a multi-agent system designed specifically for biomedical literature review. It orchestrates specialized agents for searching PubMed and other medical databases, extracting structured evidence from clinical papers, and synthesizing findings into evidence-graded summaries. Particularly useful for clinical evidence reviews, drug interaction research, and systematic reviews in medicine.
Query → Planning Agent (decomposes clinical question)
↓
Search Agent (PubMed, PMC, clinical trials)
↓
Extraction Agent (PICO, outcomes, evidence grade)
↓
Synthesis Agent (evidence summary, contradictions)
↓
Report Agent (structured review output)| Agent | Role | |-------|------| | Planner | Converts clinical question to PICO format, generates sub-queries | | Searcher | Queries PubMed, PMC, ClinicalTrials.gov | | Extractor | Extracts structured data: population, intervention, outcomes | | Synthesizer | Grades evidence, identifies consensus and contradictions | | Reporter | Generates formatted review with citations |
pythonfrom med_researcher import MedResearcher researcher = MedResearcher( llm_provider="anthropic", search_backends=["pubmed", "pmc", "clinical_trials"], ) # Clinical question result = researcher.review( question="What is the comparative efficacy of SGLT2 inhibitors " "versus GLP-1 receptor agonists for cardiovascular " "outcomes in type 2 diabetes?", max_papers=50, evidence_grading=True, ) print(result.summary) print(f"Papers analyzed: {len(result.papers)}") print(f"Evidence grade: {result.overall_grade}")
python# Automatic PICO extraction from clinical question pico = researcher.extract_pico( "Does metformin reduce cancer incidence in diabetic patients?" ) # P: patients with diabetes # I: metformin treatment # C: no metformin / other antidiabetics # O: cancer incidence # Search with PICO components result = researcher.review_pico( population="type 2 diabetes patients", intervention="metformin", comparison="placebo or other antidiabetics", outcome="cancer incidence", )
python# Evidence levels following GRADE methodology for paper in result.papers: print(f"{paper.title}") print(f" Study type: {paper.study_type}") # RCT, cohort, case-control print(f" Evidence level: {paper.evidence_level}") # High/Moderate/Low/Very Low print(f" Risk of bias: {paper.bias_risk}") print(f" Sample size: {paper.sample_size}") # Aggregate evidence summary print(f"\nOverall certainty: {result.certainty}") print(f"Recommendation strength: {result.recommendation}")
pythonresearcher = MedResearcher( search_config={ "pubmed": { "max_results": 100, "date_range": ("2020-01-01", "2025-12-31"), "article_types": ["Clinical Trial", "Meta-Analysis", "Randomized Controlled Trial"], }, "clinical_trials": { "status": ["Completed", "Active"], "phase": ["Phase 3", "Phase 4"], }, }, extraction_config={ "fields": ["population", "intervention", "comparator", "primary_outcome", "secondary_outcomes", "adverse_events", "sample_size", "follow_up"], }, )
python# Structured evidence table result.export_evidence_table("evidence_table.csv") # PRISMA flow diagram data prisma = result.prisma_flow() print(f"Identified: {prisma['identified']}") print(f"Screened: {prisma['screened']}") print(f"Included: {prisma['included']}") # Bibliography result.export_bibtex("references.bib") # Full report result.export_report("review.md", format="markdown")
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-02 | fail→pass | 11,598 | 7,636 | -34% | 1 | 1 | 0% | 2,514 | 2,842 | +13% | 0 | 0 | — |
case-03 | fail→pass | 18,441 | 5,350 | -71% | 1 | 1 | 0% | 3,758 | 2,276 | -39% | 0 | 0 | — |
case-18 | fail→pass | 10,143 | 2,409 | -76% | 1 | 1 | 0% | 1,731 | 1,432 | -17% | 0 | 0 | — |
case-19 | pass→pass | 10,549 | 8,992 | -15% | 1 | 1 | 0% | 1,888 | 2,915 | +54% | 0 | 0 | — |
case-01 | fail→pass | 20,450 | 30,187 | +48% | 1 | 1 | 0% | 4,104 | 3,287 | -20% | 0 | 0 | — |
case-04 | fail→pass | 10,031 | 3,887 | -61% | 1 | 1 | 0% | 1,632 | 1,838 | +13% | 0 | 0 | — |
case-05 | fail→pass | 10,259 | 2,938 | -71% | 1 | 1 | 0% | 1,700 | 1,732 | +2% | 0 | 0 | — |
case-06 | pass→pass | 8,116 | 3,029 | -63% | 1 | 1 | 0% | 1,326 | 1,684 | +27% | 0 | 0 | — |
case-07 | fail→pass | 20,814 | 4,332 | -79% | 1 | 1 | 0% | 1,775 | 1,901 | +7% | 0 | 0 | — |
case-08 | pass→pass | 10,290 | 3,093 | -70% | 1 | 1 | 0% | 1,690 | 1,686 | -0% | 0 | 0 | — |
case-09 | fail→pass | 5,539 | 3,445 | -38% | 1 | 1 | 0% | 947 | 1,809 | +91% | 0 | 0 | — |
case-10 | fail→pass | 14,537 | 4,457 | -69% | 1 | 1 | 0% | 2,806 | 1,956 | -30% | 0 | 0 | — |
case-11 | fail→pass | 10,036 | 3,032 | -70% | 1 | 1 | 0% | 1,686 | 1,709 | +1% | 0 | 0 | — |
case-12 | pass→pass | 8,083 | 2,043 | -75% | 1 | 1 | 0% | 1,320 | 1,448 | +10% | 0 | 0 | — |
case-13 | pass→pass | 7,390 | 3,367 | -54% | 1 | 1 | 0% | 1,125 | 1,592 | +42% | 0 | 0 | — |
case-14 | pass→pass | 10,298 | 2,538 | -75% | 1 | 1 | 0% | 1,615 | 1,620 | +0% | 0 | 0 | — |
case-15 | fail→pass | 15,780 | 2,538 | -84% | 1 | 1 | 0% | 2,497 | 1,629 | -35% | 0 | 0 | — |
case-16 | fail→pass | 11,166 | 2,410 | -78% | 1 | 1 | 0% | 1,831 | 1,563 | -15% | 0 | 0 | — |
case-17 | fail→pass | 7,932 | 3,487 | -56% | 1 | 1 | 0% | 1,291 | 1,792 | +39% | 0 | 0 | — |
case-20 | pass→pass | 8,664 | 6,435 | -26% | 1 | 1 | 0% | 1,559 | 2,352 | +51% | 0 | 0 | — |
case-21 | pass→pass | 6,794 | 7,058 | +4% | 1 | 1 | 0% | 1,098 | 2,320 | +111% | 0 | 0 | — |
case-22 | pass→pass | 8,818 | 8,981 | +2% | 1 | 1 | 0% | 1,456 | 2,580 | +77% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +59 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.