Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run and monitor PapersFlow DeepScan jobs. Use when the user wants long-running research progress, intermediate findings, final reports, or plotting from a completed run.
.claude/skills/hashgraph-online-deepscan-monitor/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-11 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-16 | ✓→✗ | ▼ Worse | 145% | 0% |
| case-13 | ✓→✓ | = Same ✓ | 4% | 0% |
| case-17 | ✓→✓ | = Same ✓ | 72% | 0% |
| case-18 | ✓→✓ | = Same ✓ | 0% | 0% |
Use this skill when the user wants Claude to manage a longer-running PapersFlow research workflow instead of a single search call.
run_deepscan to start the job.get_deepscan_live_snapshot for the best live view of:get_deepscan_status if the user only wants lightweight progress checks.finalReportAvailable is true or the run is completed, call get_deepscan_report.summarize_evidence when the user wants a cross-report summary from stored DeepScan history.run_python_plot only after you have stable report data worth plotting.get_deepscan_live_snapshot over get_deepscan_status when the user wants richer live information.When a run is still active, summarize:
Keep updates brief unless the user asks for more detail.
Use run_python_plot only for meaningful visualizations after you have stable report outputs, for example:
Do not generate plots for sparse or obviously low-quality data without saying so.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 46,790 | 16,373 | -65% | 1 | 1 | 0% | 7,151 | 895 | -87% | 0 | 0 | — |
case-02 | fail→fail | 10,308 | 16,794 | +63% | 1 | 1 | 0% | 859 | 1,128 | +31% | 0 | 0 | — |
case-03 | fail→fail | 23,965 | 14,239 | -41% | 1 | 1 | 0% | 3,688 | 1,304 | -65% | 0 | 0 | — |
case-04 | fail→fail | 50,702 | 10,889 | -79% | 1 | 1 | 0% | 8,222 | 1,344 | -84% | 0 | 0 | — |
case-05 | fail→fail | 14,076 | 14,786 | +5% | 1 | 1 | 0% | 2,101 | 914 | -56% | 0 | 0 | — |
case-06 | fail→fail | 10,498 | 12,401 | +18% | 1 | 1 | 0% | 881 | 1,402 | +59% | 0 | 0 | — |
case-07 | fail→fail | 36,280 | 4,408 | -88% | 1 | 1 | 0% | 5,120 | 831 | -84% | 0 | 0 | — |
case-08 | fail→fail | 36,727 | 20,583 | -44% | 1 | 1 | 0% | 4,360 | 907 | -79% | 0 | 0 | — |
case-09 | fail→fail | 16,850 | 13,039 | -23% | 1 | 1 | 0% | 2,798 | 1,041 | -63% | 0 | 0 | — |
case-10 | fail→fail | 13,079 | 12,000 | -8% | 1 | 1 | 0% | 2,626 | 839 | -68% | 0 | 0 | — |
case-11 | fail→pass | 15,672 | 20,462 | +31% | 1 | 1 | 0% | 1,634 | 1,865 | +14% | 0 | 0 | — |
case-12 | fail→fail | 22,247 | 6,338 | -72% | 1 | 1 | 0% | 2,870 | 1,006 | -65% | 0 | 0 | — |
case-13 | pass→pass | 17,595 | 9,735 | -45% | 1 | 1 | 0% | 1,898 | 1,968 | +4% | 0 | 0 | — |
case-14 | fail→fail | 31,266 | 10,962 | -65% | 1 | 1 | 0% | 5,883 | 878 | -85% | 0 | 0 | — |
case-15 | fail→fail | 9,017 | 7,166 | -21% | 1 | 1 | 0% | 1,266 | 1,241 | -2% | 0 | 0 | — |
case-16 | pass→fail | 5,275 | 22,288 | +323% | 1 | 1 | 0% | 853 | 2,094 | +145% | 0 | 0 | — |
case-17 | pass→pass | 13,314 | 19,299 | +45% | 1 | 1 | 0% | 1,577 | 2,711 | +72% | 0 | 0 | — |
case-18 | pass→pass | 11,838 | 9,470 | -20% | 1 | 1 | 0% | 1,411 | 1,414 | +0% | 0 | 0 | — |
case-19 | pass→pass | 13,466 | 8,002 | -41% | 1 | 1 | 0% | 1,803 | 2,145 | +19% | 0 | 0 | — |
case-20 | fail→fail | 30,522 | 11,691 | -62% | 1 | 1 | 0% | 5,326 | 914 | -83% | 0 | 0 | — |
case-21 | fail→fail | 24,789 | 9,633 | -61% | 1 | 1 | 0% | 3,580 | 961 | -73% | 0 | 0 | — |
case-22 | fail→fail | 9,084 | 9,002 | -1% | 1 | 1 | 0% | 624 | 1,022 | +64% | 0 | 0 | — |
case-23 | fail→fail | 29,125 | 16,177 | -44% | 1 | 1 | 0% | 6,171 | 1,132 | -82% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 6 counted toward the lift figure. The other 17 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of 0 percentage points is the difference between those two pass rates over the 6 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.