Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Recommend AND run open-source AI tools, agents, Claude Code / Codex skills, and MCP servers for any stage of a literature review — searching, reading, extracting, synthesizing, screening, citation-checking, and paper writing. Use when the user asks "what tool should I use to..." OR "install/run/use <tool> to ..." for research/lit-review work: automating a survey or related-work section, PDF→Markdown extraction for LLMs (MinerU/marker/docling), PRISMA / systematic review (ASReview), citation-back
.claude/skills/brycewang-stanford-literature-review-tools/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | 37% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 40% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 144% | 0% |
A curated, use-case-organized catalog of the strongest open-source AI tools for literature review — plus a launcher that actually installs and runs the top ones. Covers: end-to-end research agents, deep-research / auto-survey generators, autonomous "idea→paper" systems, citation-backed RAG over PDFs, PRISMA screening, MCP servers, Zotero/Obsidian integrations, PDF→structured extraction, citation graphs, and paper-writing / peer-review assistants.
Full source of truth (README, always current star counts): <https://github.com/brycewang-stanford/lit-review-agent-tools>
scripts/litrun.py via Bash — do not hand the user raw pip commands to copy.scripts/litrun.pyThe launcher installs each supported tool into its own venv under ~/.lit-review-tools/ (uses uv if present, else python -m venv) and reads API keys from one shared ~/.lit-review-tools/.env. Machine-readable recipes: recipes/recipes.json.
Typical flow when the user wants to use a tool:
python3 scripts/litrun.py doctor — check toolchain + which API keys are already set.python3 scripts/litrun.py info <id> — confirm what the tool needs (entry, required env).litrun.py env --set KEY=VALUE (never echo the value back in full).python3 scripts/litrun.py run <id> -- <tool args> — installs on first use, then runs. For PDF tools pass the real file path; e.g. run mineru -- -p paper.pdf -o ./out -b pipeline.litrun.py mcp <id> prints the client config block to register in Claude Code / Cursor.Commands: list [--category C] [--kind K] · info <id> · doctor · env [--set K=V] · install <id> · run <id> -- <args> · mcp <id> [--storage PATH] [--client claude|cursor] · ui <id>.
Runnable ids by kind:
mineru, marker, docling (PDF→Markdown) · paper-qa (cited Q&A) · asreview (PRISMA screening UI)arxiv-fetch (search arXiv & download PDFs, no key)gpt-researcher, storm (deep research; need API keys) · scholarly, pyalex (API clients)mcp config): arxiv-mcp-server, paper-search-mcp, zotero-mcpFor gpt-researcher and storm, litrun.py ui <id> clones the repo and launches the full web UI (GPT Researcher → FastAPI at :8000; STORM → Streamlit at :8501). These are long-running servers — launch them with a background Bash call and tell the user the URL. gpt-researcher's UI needs OPENAI_API_KEY + TAVILY_API_KEY set first (litrun writes them into the repo's .env); STORM takes its keys in the app sidebar.
For multi-tool tasks, prefer a named workflow over hand-wiring steps: litrun.py workflow list then litrun.py workflow run <id> [--input PATH] [--query "..."] [--question "..."] [--max N]. Built-ins:
pdf-to-markdown — a PDF/folder → clean Markdown (MinerU)pdf-corpus-qa — a folder of PDFs → citation-backed answer (PaperQA2)pdf-md-then-qa — convert to Markdown and answer a question over the corpustopic-to-pdfs — arXiv query → download top-N PDFs (arxiv-fetch, no key)topic-to-review — arXiv query → download PDFs → citation-backed answer (PaperQA2). The end-to-end "retrieve then review" pipeline; no MCP client needed. Needs OPENAI_API_KEY for the QA step.Add --dry-run first to show the exact resolved step commands without executing — good for confirming paths with the user before a heavy run. Workflows fail fast if a required API key is missing.
Guardrails: installs and downloads happen under the user's home and hit the network — for a heavy first install (marker/docling pull in PyTorch) say so before running. Never fabricate API keys. If a run fails, show the real error rather than claiming success. Paths in this file (scripts/…, recipes/…) are relative to this skill's directory.
reference/catalog.md. Do not guess project names or URLs; pull them from the catalog.textUse Claude Code, want end-to-end research→paper ──────────▶ academic-research-skills ⭐ Want AI to research a topic → cited report ───────────────▶ GPT Researcher / STORM Want fully autonomous "idea → submittable paper" ────────▶ AI-Scientist-v2 / AutoResearchClaw Citation-backed Q&A over a pile of PDFs ──────────────────▶ PaperQA2 Rigorous PRISMA review (thousands of abstracts) ─────────▶ ASReview / prismAId Clean Markdown from PDFs to feed an LLM ─────────────────▶ MinerU / Docling / marker Lit capabilities inside Claude / Cursor (MCP) ───────────▶ paper-search-mcp / zotero-mcp Chat with your library inside Zotero ────────────────────▶ zotero-gpt / PapersGPT Pre-submission AI peer review ───────────────────────────▶ open_reviewer / ai-peer-review
| Category | Editor's pick ⭐ | When | |---|---|---| | All-in-one research agents & skills | academic-research-skills | Claude Code user wanting research→write→review→revise, with integrity/citation gates | | Deep research & auto-survey | STORM / gpt-researcher | Topic → cited survey / report / related-work | | Autonomous science (idea→paper) | AI-Scientist(-v2) / AutoResearchClaw | Fully automated discovery: lit + hypotheses + experiments + writing | | Literature Q&A / RAG | paper-qa (PaperQA2) | Citation-backed answers over a PDF corpus | | Systematic review & screening | ASReview | Active-learning screening of thousands of abstracts (PRISMA) | | MCP servers | zotero-mcp / arxiv-mcp-server | Wire papers into Claude / Cursor / Cline | | Zotero / Obsidian integration | zotero-gpt | Chat with your library inside your reference manager | | PDF → structured extraction | MinerU / docling / marker | Turn PDFs into clean Markdown/JSON for LLMs | | Citation graphs & API clients | scholarly / pyalex | Citation-network analysis; scripting academic DBs | | Writing & peer-review assistants | open_reviewer / ai-peer-review | Draft, polish, and pre-submission review | | Awesome lists | Awesome-Auto-Research-Tools | Browse the whole landscape |
| User's need | Recommend | |---|---| | Claude Code, end-to-end research→paper | academic-research-skills (most complete, #1 in space) | | Generic "research this topic for me" agent | GPT Researcher / STORM | | Wiki/survey-style long-form with citations | STORM / Co-STORM | | Fully autonomous "idea → submittable paper" | AI-Scientist-v2 / AutoResearchClaw | | Cited Q&A over many PDFs | PaperQA / PaperQA2 | | Rigorous PRISMA systematic review | ASReview or prismAId | | PDF → clean Markdown for an LLM | MinerU / Docling / marker | | Lit capabilities in an MCP client | paper-search-mcp / zotero-mcp | | Chat with library inside Zotero | zotero-gpt / PapersGPT | | AI pre-review before submission | open_reviewer / ai-peer-review | | Just want to browse the landscape | The Awesome lists section |
local-deep-research; medical → medsci-skills / paperai; Codex instead of Claude → academic-research-skills-codex.Full catalog with every project, star count, and one-line description: reference/catalog.md.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-15 | fail→pass | 14,137 | 4,722 | -67% | 1 | 1 | 0% | 2,405 | 3,305 | +37% | 0 | 0 | — |
case-01 | fail→fail | 15,705 | 4,066 | -74% | 1 | 1 | 0% | 3,295 | 2,743 | -17% | 0 | 0 | — |
case-02 | fail→fail | 16,539 | 11,036 | -33% | 1 | 1 | 0% | 3,222 | 4,383 | +36% | 0 | 0 | — |
case-03 | fail→fail | 7,432 | 6,557 | -12% | 1 | 1 | 0% | 283 | 2,778 | +882% | 0 | 0 | — |
case-04 | pass→pass | 16,226 | 16,090 | -1% | 1 | 1 | 0% | 2,743 | 5,082 | +85% | 0 | 0 | — |
case-05 | fail→fail | 5,497 | 9,571 | +74% | 1 | 1 | 0% | 916 | 3,557 | +288% | 0 | 0 | — |
case-06 | pass→pass | 16,222 | 15,608 | -4% | 1 | 1 | 0% | 3,286 | 5,393 | +64% | 0 | 0 | — |
case-07 | fail→pass | 18,330 | 5,251 | -71% | 1 | 1 | 0% | 3,033 | 3,393 | +12% | 0 | 0 | — |
case-08 | fail→pass | 12,630 | 8,472 | -33% | 1 | 1 | 0% | 2,563 | 3,590 | +40% | 0 | 0 | — |
case-09 | fail→fail | 10,750 | 5,819 | -46% | 1 | 1 | 0% | 1,951 | 2,872 | +47% | 0 | 0 | — |
case-10 | fail→fail | 11,217 | 4,893 | -56% | 1 | 1 | 0% | 2,249 | 2,758 | +23% | 0 | 0 | — |
case-11 | fail→fail | 4,434 | 5,571 | +26% | 1 | 1 | 0% | 859 | 2,852 | +232% | 0 | 0 | — |
case-12 | fail→pass | 15,298 | 7,125 | -53% | 1 | 1 | 0% | 2,772 | 3,753 | +35% | 0 | 0 | — |
case-13 | fail→fail | 6,135 | 4,533 | -26% | 1 | 1 | 0% | 1,045 | 2,726 | +161% | 0 | 0 | — |
case-14 | fail→fail | 7,305 | 5,351 | -27% | 1 | 1 | 0% | 1,273 | 2,791 | +119% | 0 | 0 | — |
case-16 | fail→fail | 8,342 | 4,503 | -46% | 1 | 1 | 0% | 1,368 | 2,773 | +103% | 0 | 0 | — |
case-17 | pass→pass | 10,823 | 6,120 | -43% | 1 | 1 | 0% | 1,860 | 3,500 | +88% | 0 | 0 | — |
case-18 | fail→pass | 7,276 | 3,514 | -52% | 1 | 1 | 0% | 1,308 | 3,186 | +144% | 0 | 0 | — |
case-19 | fail→pass | 15,142 | 7,963 | -47% | 1 | 1 | 0% | 2,702 | 3,923 | +45% | 0 | 0 | — |
case-20 | fail→fail | 9,597 | 3,931 | -59% | 1 | 1 | 0% | 1,918 | 2,819 | +47% | 0 | 0 | — |
case-21 | fail→fail | 9,509 | 4,341 | -54% | 1 | 1 | 0% | 1,843 | 2,742 | +49% | 0 | 0 | — |
case-22 | fail→fail | 5,868 | 4,853 | -17% | 1 | 1 | 0% | 1,259 | 2,732 | +117% | 0 | 0 | — |
case-23 | fail→pass | 16,882 | 10,105 | -40% | 1 | 1 | 0% | 2,993 | 3,753 | +25% | 0 | 0 | — |
case-24 | pass→pass | 13,950 | 7,392 | -47% | 1 | 1 | 0% | 2,311 | 3,903 | +69% | 0 | 0 | — |
case-25 | fail→pass | 6,529 | 3,411 | -48% | 1 | 1 | 0% | 1,110 | 3,019 | +172% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 15 counted toward the lift figure. The other 10 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 15 comparable cases. 5 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.