Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Creates a structured annotation note in references/ with sections for research question, data, findings, and relevance. Use when documenting a paper.
.claude/skills/brycewang-stanford-literature-note/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | -26% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -60% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -36% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -50% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -44% | 0% |
Create a structured annotation note for a paper in references/.
$ARGUMENTS — a citation key from references.bib, a DOI, or a paper description (e.g., "acemoglu2001colonial" or "Acemoglu 2001 colonial origins")references.bib, use that entry's metadatareferences.bib first (offer to run the /project:cite workflow)references/ named <citation-key>.md with this structure:markdown # <Author (Year)> — <Short Title>
Citation key: <key> Full reference: <formatted reference>
## Research Question
What question does this paper address?]
## Identification Strategy
How do the authors establish causality? What is the main source of variation?]
## Data and Sample
What data do they use? What is the sample period, unit of observation, and sample size?]
## Key Findings
## Methodology Notes
Econometric methods, estimators, robustness checks worth noting]
## Relevance to This Project
How does this paper relate to the current research? What can we build on or contrast with?]
references.bib and cannot be resolved, ask the user for more details.references/<key>.md already exists, show the existing note and ask if the user wants to update it.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→fail | 4,842 | 4,627 | -4% | 1 | 1 | 0% | 1,071 | 672 | -37% | 0 | 0 | — |
case-02 | pass→fail | 9,728 | 4,717 | -52% | 1 | 1 | 0% | 1,853 | 1,460 | -21% | 0 | 0 | — |
case-03 | pass→pass | 9,166 | 8,217 | -10% | 1 | 1 | 0% | 1,929 | 1,530 | -21% | 0 | 0 | — |
case-04 | fail→fail | 15,709 | 2,816 | -82% | 1 | 1 | 0% | 2,764 | 632 | -77% | 0 | 0 | — |
case-05 | fail→fail | 12,869 | 5,563 | -57% | 1 | 1 | 0% | 1,962 | 924 | -53% | 0 | 0 | — |
case-06 | fail→fail | 11,742 | 4,754 | -60% | 1 | 1 | 0% | 2,179 | 695 | -68% | 0 | 0 | — |
case-07 | fail→fail | 12,340 | 4,939 | -60% | 1 | 1 | 0% | 2,134 | 693 | -68% | 0 | 0 | — |
case-08 | fail→fail | 8,191 | 3,590 | -56% | 1 | 1 | 0% | 1,397 | 955 | -32% | 0 | 0 | — |
case-17 | fail→fail | 6,236 | 2,206 | -65% | 1 | 1 | 0% | 815 | 713 | -13% | 0 | 0 | — |
case-09 | fail→pass | 11,688 | 5,210 | -55% | 1 | 1 | 0% | 2,018 | 1,497 | -26% | 0 | 0 | — |
case-10 | pass→pass | 5,712 | 3,082 | -46% | 1 | 1 | 0% | 1,080 | 856 | -21% | 0 | 0 | — |
case-11 | fail→pass | 11,975 | 2,563 | -79% | 1 | 1 | 0% | 1,946 | 777 | -60% | 0 | 0 | — |
case-12 | pass→fail | 12,149 | 7,626 | -37% | 1 | 1 | 0% | 2,021 | 1,733 | -14% | 0 | 0 | — |
case-13 | pass→pass | 5,664 | 3,774 | -33% | 1 | 1 | 0% | 863 | 1,075 | +25% | 0 | 0 | — |
case-14 | fail→pass | 9,434 | 3,883 | -59% | 1 | 1 | 0% | 1,737 | 1,114 | -36% | 0 | 0 | — |
case-15 | fail→pass | 9,112 | 1,578 | -83% | 1 | 1 | 0% | 1,365 | 682 | -50% | 0 | 0 | — |
case-16 | fail→pass | 9,318 | 2,465 | -74% | 1 | 1 | 0% | 1,551 | 865 | -44% | 0 | 0 | — |
case-18 | fail→pass | 5,457 | 1,224 | -78% | 1 | 1 | 0% | 893 | 625 | -30% | 0 | 0 | — |
case-19 | fail→pass | 10,566 | 1,685 | -84% | 1 | 1 | 0% | 1,802 | 683 | -62% | 0 | 0 | — |
case-20 | fail→pass | 10,527 | 2,961 | -72% | 1 | 1 | 0% | 1,788 | 913 | -49% | 0 | 0 | — |
case-21 | fail→pass | 5,842 | 1,338 | -77% | 1 | 1 | 0% | 906 | 650 | -28% | 0 | 0 | — |
case-22 | fail→pass | 8,851 | 1,909 | -78% | 1 | 1 | 0% | 1,304 | 755 | -42% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 17 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.