Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Systematic framework for evaluating scholarly and research work based on the ScholarEval methodology. This skill should be used when assessing research papers, evaluating literature reviews, scoring research methodologies, analyzing scientific writing quality, or applying structured evaluation criteria to academic work. Provides comprehensive assessment across multiple dimensions including problem formulation, literature review, methodology, data collection, analysis, results interpretation, and
.claude/skills/microck-scholar-evaluation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 190% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 54% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 65% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -4% | 0% |
Apply the ScholarEval framework to systematically evaluate scholarly and research work. This skill provides structured evaluation methodology based on peer-reviewed research assessment criteria, enabling comprehensive analysis of academic papers, research proposals, literature reviews, and scholarly writing across multiple quality dimensions.
Use this skill when:
Begin by identifying the type of scholarly work being evaluated and the evaluation scope:
Work Types:
Evaluation Scope:
Ask the user to clarify if the scope is ambiguous.
Systematically evaluate the work across the ScholarEval dimensions. For each applicable dimension, assess quality, identify strengths and weaknesses, and provide scores where appropriate.
Refer to references/evaluation_framework.md for detailed criteria and rubrics for each dimension.
Core Evaluation Dimensions:
For each evaluated dimension, provide:
Qualitative Assessment:
Quantitative Scoring (Optional): Use a 5-point scale where applicable:
To calculate aggregate scores programmatically, use scripts/calculate_scores.py.
Provide an integrated evaluation summary:
Transform evaluation findings into constructive, actionable feedback:
Feedback Structure:
Feedback Format Options:
Adjust evaluation approach based on:
Stage of Development:
Purpose and Venue:
Discipline-Specific Norms:
Detailed evaluation criteria, rubrics, and quality indicators for each ScholarEval dimension. Load this reference when conducting evaluations to access specific assessment guidelines and scoring rubrics.
Search patterns for quick access:
Python script for calculating aggregate evaluation scores from dimension-level ratings. Supports weighted averaging, threshold analysis, and score visualization.
Usage:
pythonpython scripts/calculate_scores.py --scores <dimension_scores.json> --output <report.txt>
User Request: "Evaluate this research paper on machine learning for drug discovery"
Response Process:
references/evaluation_framework.md for detailed criteriaThis skill is based on the ScholarEval framework introduced in:
Moussa, H. N., Da Silva, P. Q., Adu-Ampratwum, D., East, A., Lu, Z., Puccetti, N., Xue, M., Sun, H., Majumder, B. P., & Kumar, S. (2025). _ScholarEval: Research Idea Evaluation Grounded in Literature_. arXiv preprint arXiv:2510.16234. https://arxiv.org/abs/2510.16234
Abstract: ScholarEval is a retrieval augmented evaluation framework that assesses research ideas based on two fundamental criteria: soundness (the empirical validity of proposed methods based on existing literature) and contribution (the degree of advancement made by the idea across different dimensions relative to prior research). The framework achieves significantly higher coverage of expert-annotated evaluation points and is consistently preferred over baseline systems in terms of evaluation actionability, depth, and evidence support.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-07 | pass→pass | 10,894 | 7,799 | -28% | 1 | 1 | 0% | 1,616 | 3,139 | +94% | 0 | 0 | — |
case-08 | pass→pass | 9,270 | 5,644 | -39% | 1 | 1 | 0% | 1,373 | 2,764 | +101% | 0 | 0 | — |
case-01 | fail→fail | 7,690 | 7,311 | -5% | 1 | 1 | 0% | 1,251 | 2,975 | +138% | 0 | 0 | — |
case-02 | fail→pass | 9,340 | 13,312 | +43% | 1 | 1 | 0% | 1,419 | 4,111 | +190% | 0 | 0 | — |
case-03 | fail→pass | 10,960 | 4,444 | -59% | 1 | 1 | 0% | 1,778 | 2,744 | +54% | 0 | 0 | — |
case-04 | pass→pass | 7,001 | 6,775 | -3% | 1 | 1 | 0% | 1,085 | 2,730 | +152% | 0 | 0 | — |
case-05 | fail→pass | 13,100 | 7,367 | -44% | 1 | 1 | 0% | 1,919 | 3,172 | +65% | 0 | 0 | — |
case-06 | pass→pass | 13,962 | 10,550 | -24% | 1 | 1 | 0% | 2,173 | 3,461 | +59% | 0 | 0 | — |
case-09 | pass→pass | 12,470 | 9,997 | -20% | 1 | 1 | 0% | 1,832 | 3,417 | +87% | 0 | 0 | — |
case-10 | fail→pass | 12,453 | 5,369 | -57% | 1 | 1 | 0% | 1,938 | 2,611 | +35% | 0 | 0 | — |
case-11 | fail→pass | 15,342 | 3,215 | -79% | 1 | 1 | 0% | 2,621 | 2,512 | -4% | 0 | 0 | — |
case-12 | pass→pass | 10,248 | 4,506 | -56% | 1 | 1 | 0% | 1,589 | 2,635 | +66% | 0 | 0 | — |
case-13 | pass→pass | 17,136 | 13,464 | -21% | 1 | 1 | 0% | 2,522 | 3,894 | +54% | 0 | 0 | — |
case-14 | pass→pass | 13,466 | 9,891 | -27% | 1 | 1 | 0% | 1,993 | 3,445 | +73% | 0 | 0 | — |
case-15 | fail→fail | 10,894 | 2,460 | -77% | 1 | 1 | 0% | 1,440 | 2,237 | +55% | 0 | 0 | — |
case-16 | pass→fail | 9,929 | 2,782 | -72% | 1 | 1 | 0% | 1,515 | 2,296 | +52% | 0 | 0 | — |
case-17 | pass→pass | 14,646 | 9,743 | -33% | 1 | 1 | 0% | 2,048 | 3,237 | +58% | 0 | 0 | — |
case-18 | fail→pass | 26,424 | 10,289 | -61% | 1 | 1 | 0% | 2,227 | 3,442 | +55% | 0 | 0 | — |
case-19 | pass→pass | 12,659 | 7,354 | -42% | 1 | 1 | 0% | 1,873 | 3,086 | +65% | 0 | 0 | — |
case-20 | pass→pass | 22,828 | 24,632 | +8% | 1 | 1 | 0% | 4,212 | 6,638 | +58% | 0 | 0 | — |
case-21 | fail→fail | 3,711 | 4,524 | +22% | 1 | 1 | 0% | 552 | 2,564 | +364% | 0 | 0 | — |
case-22 | pass→pass | 18,052 | 23,341 | +29% | 1 | 1 | 0% | 2,571 | 5,129 | +99% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +23 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.