Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build a focused literature and citation briefing from PapersFlow. Use when the user wants paper search, citation verification, related-paper discovery, or citation graph exploration.
.claude/skills/hashgraph-online-research-briefing/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | -48% | 0% |
| case-08 | ✓→✗ | ▼ Worse | 26% | 0% |
| case-09 | ✓→✗ | ▼ Worse | -83% | 0% |
| case-13 | ✓→✗ | ▼ Worse | -55% | 0% |
| case-19 | ✓→✗ | ▼ Worse | -63% | 0% |
Use this skill when a user wants a high-signal research briefing grounded in the hosted papersflow-mcp server.
search_literature for broad discovery.verify_citation when the user gives a citation string, DOI, URL, PubMed ID, arXiv ID, or uncertain bibliographic reference.find_related_papers when the user wants nearby work around a seed paper.get_citation_graph when the user wants a graph view with references, incoming citations, and optionally similar papers.get_paper_neighbors when the user wants a one-hop grouped view instead of a graph-first view.expand_citation_graph only after you already have seed node ids from a previous graph result.fetch when the user wants a richer single-paper record after search or graph exploration.search_literatureUse for:
verify_citationUse for:
get_citation_graphUse for:
Prefer this when the user explicitly asks for a graph, network, map, or influence chain.
get_paper_neighborsUse for:
Prefer this over the graph tool when the user cares more about grouped lists than graph structure.
expand_citation_graphUse only when:
Do not guess node ids.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 15,868 | 19,164 | +21% | 1 | 1 | 0% | 2,827 | 896 | -68% | 0 | 0 | — |
case-02 | fail→fail | 13,770 | 10,527 | -24% | 1 | 1 | 0% | 1,672 | 904 | -46% | 0 | 0 | — |
case-03 | fail→fail | 15,434 | 29,627 | +92% | 1 | 1 | 0% | 3,105 | 1,039 | -67% | 0 | 0 | — |
case-04 | fail→fail | 15,917 | 8,408 | -47% | 1 | 1 | 0% | 3,224 | 1,130 | -65% | 0 | 0 | — |
case-14 | fail→fail | 12,463 | 10,802 | -13% | 1 | 1 | 0% | 1,433 | 951 | -34% | 0 | 0 | — |
case-05 | fail→pass | 24,769 | 9,449 | -62% | 1 | 1 | 0% | 3,854 | 2,005 | -48% | 0 | 0 | — |
case-06 | fail→fail | 20,198 | 7,408 | -63% | 1 | 1 | 0% | 3,040 | 936 | -69% | 0 | 0 | — |
case-07 | fail→fail | 8,376 | 13,287 | +59% | 1 | 1 | 0% | 1,643 | 1,317 | -20% | 0 | 0 | — |
case-08 | pass→fail | 20,783 | 18,206 | -12% | 1 | 1 | 0% | 2,126 | 2,683 | +26% | 0 | 0 | — |
case-09 | pass→fail | 30,521 | 16,690 | -45% | 1 | 1 | 0% | 4,975 | 821 | -83% | 0 | 0 | — |
case-10 | fail→fail | 7,413 | 10,614 | +43% | 1 | 1 | 0% | 1,463 | 850 | -42% | 0 | 0 | — |
case-11 | fail→fail | 21,640 | 12,217 | -44% | 1 | 1 | 0% | 2,838 | 985 | -65% | 0 | 0 | — |
case-12 | fail→fail | 18,924 | 5,568 | -71% | 1 | 1 | 0% | 2,931 | 915 | -69% | 0 | 0 | — |
case-13 | pass→fail | 39,971 | 7,212 | -82% | 1 | 1 | 0% | 2,317 | 1,038 | -55% | 0 | 0 | — |
case-15 | fail→fail | 4,445 | 9,733 | +119% | 1 | 1 | 0% | 388 | 1,226 | +216% | 0 | 0 | — |
case-16 | fail→fail | 18,422 | 12,705 | -31% | 1 | 1 | 0% | 2,312 | 968 | -58% | 0 | 0 | — |
case-17 | fail→fail | 9,963 | 10,772 | +8% | 1 | 1 | 0% | 1,463 | 920 | -37% | 0 | 0 | — |
case-18 | fail→fail | 17,521 | 12,910 | -26% | 1 | 1 | 0% | 2,659 | 1,003 | -62% | 0 | 0 | — |
case-19 | pass→fail | 20,073 | 16,169 | -19% | 1 | 1 | 0% | 2,434 | 900 | -63% | 0 | 0 | — |
case-20 | pass→fail | 36,532 | 13,237 | -64% | 1 | 1 | 0% | 5,240 | 936 | -82% | 0 | 0 | — |
case-21 | pass→pass | 4,538 | 10,073 | +122% | 1 | 1 | 0% | 917 | 1,488 | +62% | 0 | 0 | — |
case-22 | pass→fail | 18,215 | 13,356 | -27% | 1 | 1 | 0% | 2,644 | 1,112 | -58% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 4 counted toward the lift figure. The other 18 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -23 percentage points is the difference between those two pass rates over the 4 comparable cases. 9 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.