Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Searches academic literature via arXiv, Semantic Scholar, and open-access PDFs. Use when building literature reviews or finding formal research on a topic.
.claude/skills/athola-papers/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-18 | ✗→✓ | ▲ Improved | -46% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -45% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 109% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -37% | 0% |
| case-19 | ✗→✓ | ▲ Improved | -48% | 0% |
/tome:discourse)/tome:code-search)Search arXiv, Semantic Scholar, and open-access sources.
After acquiring a paper URL or local file path, convert the PDF to markdown for better extraction quality.
Apply the leyline:document-conversion protocol:
convert_to_markdowntool with the PDF URL or file:// path. This produces structured markdown preserving tables, equations, figures, and section hierarchy.
pages: "1-20" for the first chunkpages: "21-40" for longer papersFrom the converted markdown, extract:
When a paper is paywalled and no open version exists:
(arXiv first, then Semantic Scholar, then Unpaywall/CORE/PubMed)
leyline:document-conversionprotocol: markitdown MCP attempted first, Read-tool fallback used if markitdown is unavailable
and an abstract or key findings summary
guidance (library, DeepDyve, ILL, author contact) rather than returning an empty result
rather than generating fabricated paper entries
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-18 | fail→pass | 16,007 | 4,696 | -71% | 1 | 1 | 0% | 2,550 | 1,388 | -46% | 0 | 0 | — |
case-04 | pass→pass | 15,632 | 6,803 | -56% | 1 | 1 | 0% | 2,455 | 1,758 | -28% | 0 | 0 | — |
case-09 | fail→pass | 11,148 | 2,275 | -80% | 1 | 1 | 0% | 1,832 | 1,012 | -45% | 0 | 0 | — |
case-01 | fail→fail | 21,884 | 10,880 | -50% | 1 | 1 | 0% | 3,856 | 1,190 | -69% | 0 | 0 | — |
case-02 | fail→pass | 18,771 | 33,955 | +81% | 1 | 1 | 0% | 3,071 | 6,409 | +109% | 0 | 0 | — |
case-03 | fail→fail | 20,171 | 8,777 | -56% | 1 | 1 | 0% | 3,802 | 841 | -78% | 0 | 0 | — |
case-05 | pass→fail | 16,849 | 14,178 | -16% | 1 | 1 | 0% | 2,962 | 3,104 | +5% | 0 | 0 | — |
case-06 | fail→fail | 16,308 | 20,005 | +23% | 1 | 1 | 0% | 2,363 | 3,519 | +49% | 0 | 0 | — |
case-07 | pass→pass | 10,851 | 4,570 | -58% | 1 | 1 | 0% | 1,647 | 1,276 | -23% | 0 | 0 | — |
case-08 | fail→fail | 10,420 | 2,998 | -71% | 1 | 1 | 0% | 1,475 | 1,124 | -24% | 0 | 0 | — |
case-10 | fail→fail | 6,394 | 1,946 | -70% | 1 | 1 | 0% | 974 | 900 | -8% | 0 | 0 | — |
case-11 | pass→pass | 13,396 | 6,787 | -49% | 1 | 1 | 0% | 2,057 | 1,603 | -22% | 0 | 0 | — |
case-12 | pass→pass | 9,196 | 4,891 | -47% | 1 | 1 | 0% | 1,262 | 1,296 | +3% | 0 | 0 | — |
case-13 | fail→fail | 11,749 | 6,391 | -46% | 1 | 1 | 0% | 1,922 | 1,609 | -16% | 0 | 0 | — |
case-14 | fail→pass | 14,999 | 5,282 | -65% | 1 | 1 | 0% | 2,235 | 1,409 | -37% | 0 | 0 | — |
case-15 | pass→pass | 5,804 | 1,828 | -69% | 1 | 1 | 0% | 900 | 842 | -6% | 0 | 0 | — |
case-16 | fail→fail | 16,185 | 16,229 | +0% | 1 | 1 | 0% | 2,610 | 3,313 | +27% | 0 | 0 | — |
case-17 | fail→fail | 13,355 | 10,099 | -24% | 1 | 1 | 0% | 2,081 | 2,205 | +6% | 0 | 0 | — |
case-19 | fail→pass | 12,846 | 2,738 | -79% | 1 | 1 | 0% | 1,917 | 997 | -48% | 0 | 0 | — |
case-20 | pass→pass | 6,834 | 3,598 | -47% | 1 | 1 | 0% | 982 | 1,179 | +20% | 0 | 0 | — |
case-21 | fail→fail | 14,440 | 2,556 | -82% | 1 | 1 | 0% | 2,185 | 881 | -60% | 0 | 0 | — |
case-22 | pass→pass | 16,646 | 6,942 | -58% | 1 | 1 | 0% | 2,370 | 1,673 | -29% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +18 percentage points is the difference between those two pass rates over the 20 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.