Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Automate systematic literature reviews with LatteReview AI agents
.claude/skills/brycewang-stanford-latte-review-guide/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 11% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -18% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -6% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -21% | 0% |
LatteReview is a low-code Python package that uses AI agents to automate systematic literature reviews. It handles title/abstract screening, full-text assessment, data extraction, and PRISMA-compliant reporting — tasks that typically consume hundreds of researcher-hours. Supports multiple LLM backends (Anthropic, OpenAI, local models).
bashpip install lattereview
pythonfrom lattereview import ReviewProject # Create a new review project project = ReviewProject( name="ML in Medical Imaging Review", research_question="What deep learning architectures are used for " "medical image segmentation?", inclusion_criteria=[ "Uses deep learning for medical image segmentation", "Published in peer-reviewed venue", "Reports quantitative evaluation metrics", ], exclusion_criteria=[ "Review/survey articles", "Non-English publications", "Conference abstracts only", ], )
python# Import from various sources project.import_papers("scopus_export.csv", source="scopus") project.import_papers("pubmed_export.csv", source="pubmed") # Or from a DataFrame import pandas as pd df = pd.read_csv("papers.csv") project.import_from_dataframe(df, title_col="title", abstract_col="abstract", year_col="year", ) print(f"Imported {project.total_papers} papers")
pythonfrom lattereview.agents import ScreeningAgent # Configure screening agent screener = ScreeningAgent( llm_provider="anthropic", model="claude-sonnet-4-20250514", criteria=project.inclusion_criteria, exclusion=project.exclusion_criteria, ) # Title/abstract screening results = screener.screen( project.papers, mode="title_abstract", confidence_threshold=0.7, ) # Results include: decision, confidence, reasoning for paper in results[:3]: print(f"{paper.title}") print(f" Decision: {paper.decision} " f"(confidence: {paper.confidence:.2f})") print(f" Reason: {paper.reasoning}")
pythonfrom lattereview.agents import ExtractionAgent extractor = ExtractionAgent( llm_provider="anthropic", fields={ "architecture": "Deep learning architecture used", "dataset": "Medical imaging dataset", "modality": "Imaging modality (CT, MRI, X-ray, etc.)", "dice_score": "Best Dice similarity coefficient reported", "sample_size": "Number of images/patients", }, ) extracted = extractor.extract(project.included_papers) # Export structured data extracted.to_csv("extracted_data.csv")
python# PRISMA flow diagram project.generate_prisma_diagram("prisma.png") # Summary statistics summary = project.summarize() print(f"Screened: {summary['screened']}") print(f"Included: {summary['included']}") print(f"Excluded: {summary['excluded']}")
python# Use different LLM providers screener = ScreeningAgent( llm_provider="openai", model="gpt-4o", ) # Local models via Ollama screener = ScreeningAgent( llm_provider="ollama", model="llama3", base_url="http://localhost:11434", )
python# Simulate dual-reviewer screening for reliability results = screener.dual_screen( project.papers, models=["claude-sonnet-4-20250514", "gpt-4o"], agreement_threshold=0.8, ) # Papers with disagreement flagged for human review conflicts = [p for p in results if p.agreement < 0.8] print(f"{len(conflicts)} papers need human adjudication")
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 13,212 | 9,601 | -27% | 1 | 1 | 0% | 2,599 | 2,885 | +11% | 0 | 0 | — |
case-02 | fail→pass | 15,845 | 8,581 | -46% | 1 | 1 | 0% | 3,447 | 2,815 | -18% | 0 | 0 | — |
case-03 | fail→pass | 42,526 | 13,235 | -69% | 1 | 1 | 0% | 3,996 | 3,757 | -6% | 0 | 0 | — |
case-04 | pass→pass | 23,871 | 26,199 | +10% | 1 | 1 | 0% | 4,635 | 6,658 | +44% | 0 | 0 | — |
case-05 | pass→pass | 24,443 | 16,648 | -32% | 1 | 1 | 0% | 3,658 | 4,893 | +34% | 0 | 0 | — |
case-06 | pass→pass | 33,381 | 31,973 | -4% | 1 | 1 | 0% | 4,759 | 5,512 | +16% | 0 | 0 | — |
case-07 | fail→pass | 13,117 | 2,914 | -78% | 1 | 1 | 0% | 2,155 | 1,550 | -28% | 0 | 0 | — |
case-08 | fail→pass | 14,818 | 6,079 | -59% | 1 | 1 | 0% | 2,667 | 2,108 | -21% | 0 | 0 | — |
case-09 | fail→pass | 16,558 | 7,059 | -57% | 1 | 1 | 0% | 2,826 | 2,596 | -8% | 0 | 0 | — |
case-10 | fail→pass | 20,697 | 2,594 | -87% | 1 | 1 | 0% | 3,470 | 1,559 | -55% | 0 | 0 | — |
case-11 | fail→pass | 9,288 | 2,733 | -71% | 1 | 1 | 0% | 1,635 | 1,638 | +0% | 0 | 0 | — |
case-12 | fail→pass | 13,466 | 3,526 | -74% | 1 | 1 | 0% | 2,199 | 1,873 | -15% | 0 | 0 | — |
case-13 | fail→pass | 14,302 | 6,093 | -57% | 1 | 1 | 0% | 2,027 | 2,088 | +3% | 0 | 0 | — |
case-14 | fail→pass | 20,185 | 3,826 | -81% | 1 | 1 | 0% | 3,022 | 1,903 | -37% | 0 | 0 | — |
case-15 | fail→pass | 7,754 | 2,369 | -69% | 1 | 1 | 0% | 1,446 | 1,482 | +2% | 0 | 0 | — |
case-16 | fail→pass | 10,694 | 1,990 | -81% | 1 | 1 | 0% | 1,611 | 1,505 | -7% | 0 | 0 | — |
case-17 | fail→pass | 12,896 | 4,277 | -67% | 1 | 1 | 0% | 2,023 | 1,790 | -12% | 0 | 0 | — |
case-18 | fail→pass | 11,545 | 3,259 | -72% | 1 | 1 | 0% | 1,781 | 1,611 | -10% | 0 | 0 | — |
case-19 | pass→pass | 7,402 | 2,181 | -71% | 1 | 1 | 0% | 1,188 | 1,434 | +21% | 0 | 0 | — |
case-20 | fail→pass | 4,636 | 1,374 | -70% | 1 | 1 | 0% | 680 | 1,288 | +89% | 0 | 0 | — |
case-21 | fail→pass | 11,722 | 3,367 | -71% | 1 | 1 | 0% | 2,136 | 1,806 | -15% | 0 | 0 | — |
case-22 | fail→pass | 6,189 | 1,926 | -69% | 1 | 1 | 0% | 964 | 1,469 | +52% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +82 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.