Install any skill in seconds. Free to start, no credit card required.
Get Started Free →論文、提案書、文献レビュー、方法論セクション、証拠の質、引用サポート、研究論文フィードバックのための構造化された学術的作業評価。
.claude/skills/affaan-m-scholar-evaluation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 228% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 58% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 281% | 0% |
| case-02 | ✓→✓ | = Same ✓ | -30% | 0% |
このスキルを使用して、再現可能なルーブリックで学術的または科学的な作業を評価します。
まず成果物を特定します:
次に範囲を選択します:
該当する各次元を1から5でスコアリング:
適用されない次元には N/A を使用します。
markdown# Scholar Evaluation: <成果物> ## 総合評価 - 総合スコア: <1-5 または N/A> - 信頼度: <高 | 中 | 低> - サマリー: <3-5 文> ## 次元スコア | 次元 | スコア | 証拠 | 改訂優先度 | | --- | ---: | --- | --- | | 問題と質問 | | | | | 文献とコンテキスト | | | | | 方法論 | | | | | データと証拠 | | | | | 分析 | | | | | 結果と解釈 | | | | | 限界 | | | | | 文章と構造 | | | | | 引用 | | | | ## 重大な問題 ## 推奨される改訂 ## 必要な証拠確認
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 6,735 | 5,968 | -11% | 1 | 1 | 0% | 1,112 | 2,312 | +108% | 0 | 0 | — |
case-02 | pass→pass | 27,080 | 5,502 | -80% | 1 | 1 | 0% | 3,267 | 2,297 | -30% | 0 | 0 | — |
case-03 | fail→fail | 10,779 | 5,931 | -45% | 1 | 1 | 0% | 1,730 | 2,226 | +29% | 0 | 0 | — |
case-04 | fail→fail | 2,695 | 5,575 | +107% | 1 | 1 | 0% | 381 | 2,126 | +458% | 0 | 0 | — |
case-05 | pass→pass | 10,389 | 11,447 | +10% | 1 | 1 | 0% | 1,649 | 3,117 | +89% | 0 | 0 | — |
case-06 | fail→fail | 3,115 | 3,614 | +16% | 1 | 1 | 0% | 370 | 1,870 | +405% | 0 | 0 | — |
case-07 | fail→pass | 29,686 | 15,165 | -49% | 1 | 1 | 0% | 2,351 | 3,552 | +51% | 0 | 0 | — |
case-08 | pass→pass | 14,105 | 9,489 | -33% | 1 | 1 | 0% | 1,908 | 2,553 | +34% | 0 | 0 | — |
case-09 | fail→fail | 8,483 | 6,172 | -27% | 1 | 1 | 0% | 1,206 | 2,184 | +81% | 0 | 0 | — |
case-10 | fail→fail | 9,708 | 37,327 | +284% | 1 | 1 | 0% | 1,499 | 2,420 | +61% | 0 | 0 | — |
case-11 | pass→pass | 14,251 | 15,654 | +10% | 1 | 1 | 0% | 2,172 | 3,644 | +68% | 0 | 0 | — |
case-12 | fail→fail | 5,780 | 4,261 | -26% | 1 | 1 | 0% | 803 | 1,845 | +130% | 0 | 0 | — |
case-13 | fail→fail | 10,361 | 5,707 | -45% | 1 | 1 | 0% | 1,557 | 2,163 | +39% | 0 | 0 | — |
case-14 | fail→fail | 6,881 | 5,682 | -17% | 1 | 1 | 0% | 1,050 | 2,195 | +109% | 0 | 0 | — |
case-15 | fail→fail | 11,655 | 8,874 | -24% | 1 | 1 | 0% | 1,652 | 2,656 | +61% | 0 | 0 | — |
case-16 | fail→fail | 7,385 | 3,986 | -46% | 1 | 1 | 0% | 1,063 | 1,873 | +76% | 0 | 0 | — |
case-17 | fail→fail | 11,008 | 8,581 | -22% | 1 | 1 | 0% | 1,608 | 2,593 | +61% | 0 | 0 | — |
case-18 | fail→fail | 10,820 | 4,559 | -58% | 1 | 1 | 0% | 1,597 | 1,915 | +20% | 0 | 0 | — |
case-19 | fail→fail | 6,242 | 3,256 | -48% | 1 | 1 | 0% | 1,049 | 1,818 | +73% | 0 | 0 | — |
case-20 | fail→pass | 7,136 | 16,034 | +125% | 1 | 1 | 0% | 1,146 | 3,760 | +228% | 0 | 0 | — |
case-21 | fail→pass | 28,965 | 15,098 | -48% | 1 | 1 | 0% | 2,332 | 3,677 | +58% | 0 | 0 | — |
case-22 | fail→pass | 6,480 | 14,339 | +121% | 1 | 1 | 0% | 862 | 3,285 | +281% | 0 | 0 | — |
case-23 | fail→fail | 8,040 | 6,015 | -25% | 1 | 1 | 0% | 1,214 | 2,104 | +73% | 0 | 0 | — |
case-24 | fail→fail | 11,381 | 3,520 | -69% | 1 | 1 | 0% | 1,547 | 1,869 | +21% | 0 | 0 | — |
case-25 | fail→fail | 14,553 | 3,558 | -76% | 1 | 1 | 0% | 1,389 | 1,740 | +25% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted. The headline lift of 0 percentage points is the difference between those two pass rates over the 25 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.