Install any skill in seconds. Free to start, no credit card required.
Get Started Free →AI-powered paper summarization plugin for Zotero
.claude/skills/brycewang-stanford-zotero-ai-butler-guide/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 33% | 0% |
| case-18 | ✓→✗ | ▼ Worse | 19% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 34% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 5% | 0% |
Zotero AI Butler is a Zotero plugin that uses LLMs to summarize, analyze, and annotate academic papers directly within Zotero. It can generate structured summaries, extract key findings, compare papers, and answer questions about documents — all without leaving the reference manager. Supports multiple LLM backends (OpenAI, Claude, local models).
bash# Download .xpi from GitHub releases # Zotero 7: Tools → Add-ons → Install Add-on From File
markdown### LLM Backend Setup (Preferences → AI Butler) **Option 1: OpenAI** - Provider: OpenAI - Model: gpt-4o - Set environment variable for credentials **Option 2: Anthropic** - Provider: Anthropic - Model: claude-sonnet-4-20250514 **Option 3: Local (Ollama)** - Provider: Ollama - Endpoint: http://localhost:11434 - Model: llama3.1 **Option 4: Custom API** - Provider: Custom - Endpoint: your-api-url - Compatible with OpenAI API format
markdown### Usage 1. Select paper in Zotero 2. Right-click → AI Butler → Summarize 3. Summary added as Zotero note ### Summary Templates - **Quick Summary** (1 paragraph): Core contribution + method + result - **Structured Summary**: Background / Method / Results / Limitations - **Executive Brief**: Who should read this and why - **Technical Deep-Dive**: Detailed methodology and math
markdown### Extract structured information: - **Research question**: What problem does this paper address? - **Methodology**: What approach do the authors use? - **Key results**: What are the main findings? - **Contributions**: What is novel about this work? - **Limitations**: What are the acknowledged limitations? - **Future work**: What directions do the authors suggest?
markdown### Compare multiple papers: 1. Select 2+ papers in Zotero 2. Right-click → AI Butler → Compare Papers 3. Generates comparison table: - Shared and unique contributions - Methodological differences - Performance comparison (if applicable) - Complementary insights
markdown### Ask questions about papers: 1. Open paper in Zotero reader 2. AI Butler sidebar → Ask a question 3. Answers grounded in paper content with page references Example questions: - "What loss function do they use?" - "How does this compare to prior work?" - "What are the hyperparameters?" - "Explain equation 3 in simpler terms"
markdown### Summarize multiple papers: 1. Select papers (or entire collection) 2. Right-click → AI Butler → Batch Summarize 3. Progress bar shows completion 4. Each paper gets a summary note attached ### Reading List Generation: 1. Select collection 2. AI Butler → Generate Reading Order 3. Suggests optimal reading sequence based on: - Citation relationships - Conceptual dependencies - Publication chronology
markdown### Create custom analysis prompts: # In AI Butler preferences → Custom Prompts Prompt: "Systematic Review Extraction" Template: | Extract the following from this paper: 1. Study design (RCT, cohort, etc.) 2. Sample size 3. Primary outcome 4. Effect size with CI 5. Risk of bias indicators Format as structured JSON.
markdown### Combined Plugin Workflow 1. **Zotero Connector** → Import paper 2. **Zotero Sci-Hub** → Fetch PDF 3. **AI Butler** → Generate summary note 4. **Zotero Actions Tags** → Auto-tag based on summary 5. **Notero** → Sync to Notion with summary 6. **Better BibTeX** → Export citations for writing
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 16,554 | 15,346 | -7% | 1 | 1 | 0% | 2,701 | 2,913 | +8% | 0 | 0 | — |
case-02 | fail→pass | 15,929 | 11,508 | -28% | 1 | 1 | 0% | 2,373 | 2,903 | +22% | 0 | 0 | — |
case-03 | fail→fail | 19,511 | 15,194 | -22% | 1 | 1 | 0% | 2,724 | 3,589 | +32% | 0 | 0 | — |
case-04 | pass→pass | 10,017 | 6,468 | -35% | 1 | 1 | 0% | 1,638 | 2,188 | +34% | 0 | 0 | — |
case-05 | pass→pass | 14,638 | 9,033 | -38% | 1 | 1 | 0% | 2,336 | 2,448 | +5% | 0 | 0 | — |
case-06 | pass→pass | 8,217 | 3,814 | -54% | 1 | 1 | 0% | 1,251 | 1,648 | +32% | 0 | 0 | — |
case-11 | fail→fail | 11,420 | 2,706 | -76% | 1 | 1 | 0% | 1,656 | 1,539 | -7% | 0 | 0 | — |
case-07 | fail→pass | 7,424 | 2,433 | -67% | 1 | 1 | 0% | 1,079 | 1,435 | +33% | 0 | 0 | — |
case-08 | pass→pass | 10,697 | 4,765 | -55% | 1 | 1 | 0% | 1,667 | 1,807 | +8% | 0 | 0 | — |
case-09 | pass→pass | 4,733 | 3,293 | -30% | 1 | 1 | 0% | 631 | 1,598 | +153% | 0 | 0 | — |
case-10 | pass→pass | 7,112 | 3,494 | -51% | 1 | 1 | 0% | 1,084 | 1,637 | +51% | 0 | 0 | — |
case-12 | pass→pass | 9,091 | 2,714 | -70% | 1 | 1 | 0% | 1,501 | 1,513 | +1% | 0 | 0 | — |
case-13 | pass→pass | 13,240 | 5,285 | -60% | 1 | 1 | 0% | 1,876 | 1,921 | +2% | 0 | 0 | — |
case-14 | pass→pass | 13,633 | 7,371 | -46% | 1 | 1 | 0% | 1,851 | 2,184 | +18% | 0 | 0 | — |
case-15 | pass→pass | 16,495 | 3,480 | -79% | 1 | 1 | 0% | 2,618 | 1,645 | -37% | 0 | 0 | — |
case-16 | pass→pass | 9,246 | 2,745 | -70% | 1 | 1 | 0% | 1,536 | 1,559 | +1% | 0 | 0 | — |
case-17 | pass→pass | 9,845 | 6,343 | -36% | 1 | 1 | 0% | 1,464 | 1,896 | +30% | 0 | 0 | — |
case-18 | pass→fail | 11,527 | 6,719 | -42% | 1 | 1 | 0% | 1,764 | 2,096 | +19% | 0 | 0 | — |
case-19 | pass→pass | 14,697 | 5,062 | -66% | 1 | 1 | 0% | 1,881 | 1,861 | -1% | 0 | 0 | — |
case-20 | pass→pass | 12,391 | 10,368 | -16% | 1 | 1 | 0% | 2,026 | 2,744 | +35% | 0 | 0 | — |
case-21 | pass→pass | 6,800 | 7,324 | +8% | 1 | 1 | 0% | 1,140 | 2,331 | +104% | 0 | 0 | — |
case-22 | pass→pass | 11,410 | 7,720 | -32% | 1 | 1 | 0% | 1,540 | 2,317 | +50% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +5 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.