Install any skill in seconds. Free to start, no credit card required.
Get Started Free →This skill should be used when evaluating the quality of book chapters, lessons, or educational content. It provides a systematic 6-category rubric with weighted scoring (Technical Accuracy 30%, Pedagogical Effectiveness 25%, Writing Quality 20%, Structure & Organization 15%, AI-First Teaching 10%, Constitution Compliance Pass/Fail) and multi-tier assessment (Excellent/Good/Needs Work/Insufficient). Use this during iterative drafting, after content completion, on-demand review requests, or befor
.claude/skills/aiskillstore-content-evaluation-framework/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 74% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 115% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 111% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 89% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 118% | 0% |
This skill provides a comprehensive, systematic rubric for evaluating educational book chapters and lessons with quantifiable quality standards.
Constitution Alignment: v4.0.1 emphasizing:
Evaluate educational content across 6 weighted categories to ensure:
Invoke this evaluation framework at multiple checkpoints:
Multi-Tier Assessment:
The evaluation uses 6 categories with the following weights:
| Category | Weight | Focus Area | |----------|--------|------------| | Technical Accuracy | 30% | Code correctness, type hints, explanations, examples work as stated | | Pedagogical Effectiveness | 25% | Show-then-explain pattern, progressive complexity, quality exercises | | Writing Quality | 20% | Readability (Flesch-Kincaid 8-10), voice, clarity, grade-level appropriateness | | Structure & Organization | 15% | Learning objectives met, logical flow, appropriate length, transitions | | AI-First Teaching | 10% | Co-learning partnership demonstrated, Three Roles Framework shown, Nine Pillars aligned, Specs-As-Syntax emphasized | | Constitution Compliance | Pass/Fail | Must pass all non-negotiable constitutional requirements including Nine Pillars alignment (gate) |
Total Weighted Score Calculation:
Final Score = (Technical × 0.30) + (Pedagogical × 0.25) + (Writing × 0.20) +
(Structure × 0.15) + (AI-First × 0.10)Constitution Compliance: Must achieve "Pass" status. If "Fail," content cannot proceed regardless of weighted score.
Before evaluation, gather:
specs/<feature>/.specify/memory/constitution.md).claude/output-styles/lesson.md or similar)Read the detailed tier criteria for each category:
Read: references/rubric-details.mdThis file contains specific criteria defining Excellent/Good/Needs Work/Insufficient for each of the 6 categories.
Constitution compliance is a gate - if content fails constitutional requirements, it cannot proceed.
Use the constitution checklist:
Read: references/constitution-checklist.mdAssess all non-negotiable principles and requirements. Mark as Pass or Fail with specific violations noted.
If Constitution Compliance = Fail: Stop evaluation and report violations immediately. Content must be revised before proceeding.
If Constitution Compliance = Pass: Continue to weighted category evaluation.
For each of the 5 weighted categories (Technical Accuracy, Pedagogical Effectiveness, Writing Quality, Structure & Organization, AI-First Teaching):
rubric-details.md for that categoryApply the weighted formula:
Final Score = (Technical × 0.30) + (Pedagogical × 0.25) + (Writing × 0.20) +
(Structure × 0.15) + (AI-First × 0.10)Convert tier scores to numeric values:
(Or use specific numeric score within tier range if warranted)
Use the structured evaluation template:
Read: references/evaluation-template.mdComplete all sections:
Present evaluation report with:
Context: Writer requests feedback on partial draft Approach:
Context: Writer believes content is complete and ready for validation Approach:
Context: Content enters SDD Validate phase Approach:
Context: User asks "How's this looking?" for specific section Approach:
This skill includes detailed reference materials:
references/rubric-details.md - Comprehensive tier criteria for all 6 categories with specific indicatorsreferences/constitution-checklist.md - Pass/Fail checklist for constitutional compliance evaluationreferences/evaluation-template.md - Structured template for consistent evaluation reportsLoad these references as needed during evaluation to ensure consistency and thoroughness.
User Request: "Please evaluate this lesson draft: apps/learn-app/docs/chapter-3/lesson-2.md"
Evaluation Process:
apps/learn-app/docs/chapter-3/lesson-2.mdreferences/constitution-checklist.mdreferences/rubric-details.mdreferences/evaluation-template.mdUse this skill to maintain consistent, objective, evidence-based quality standards for all educational content.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-21 | fail→fail | 27,190 | 19,068 | -30% | 1 | 1 | 0% | 4,736 | 5,957 | +26% | 0 | 0 | — |
case-06 | fail→pass | 14,783 | 7,915 | -46% | 1 | 1 | 0% | 2,324 | 4,036 | +74% | 0 | 0 | — |
case-12 | fail→pass | 13,213 | 9,095 | -31% | 1 | 1 | 0% | 1,998 | 4,294 | +115% | 0 | 0 | — |
case-13 | fail→pass | 12,323 | 7,767 | -37% | 1 | 1 | 0% | 1,950 | 4,112 | +111% | 0 | 0 | — |
case-14 | fail→pass | 12,065 | 5,610 | -54% | 1 | 1 | 0% | 1,937 | 3,661 | +89% | 0 | 0 | — |
case-20 | fail→fail | 19,872 | 22,466 | +13% | 1 | 1 | 0% | 3,810 | 6,896 | +81% | 0 | 0 | — |
case-01 | fail→fail | 5,535 | 9,865 | +78% | 1 | 1 | 0% | 910 | 3,035 | +234% | 0 | 0 | — |
case-02 | fail→fail | 5,492 | 5,730 | +4% | 1 | 1 | 0% | 961 | 3,050 | +217% | 0 | 0 | — |
case-03 | fail→fail | 16,450 | 3,050 | -81% | 1 | 1 | 0% | 2,810 | 2,998 | +7% | 0 | 0 | — |
case-04 | pass→pass | 7,296 | 5,579 | -24% | 1 | 1 | 0% | 1,113 | 3,547 | +219% | 0 | 0 | — |
case-05 | fail→pass | 12,278 | 9,632 | -22% | 1 | 1 | 0% | 1,975 | 4,307 | +118% | 0 | 0 | — |
case-07 | fail→pass | 6,817 | 2,749 | -60% | 1 | 1 | 0% | 1,058 | 3,075 | +191% | 0 | 0 | — |
case-08 | pass→pass | 4,624 | 4,251 | -8% | 1 | 1 | 0% | 753 | 3,375 | +348% | 0 | 0 | — |
case-09 | pass→pass | 10,464 | 6,901 | -34% | 1 | 1 | 0% | 1,665 | 3,813 | +129% | 0 | 0 | — |
case-10 | pass→pass | 11,721 | 3,046 | -74% | 1 | 1 | 0% | 1,769 | 3,105 | +76% | 0 | 0 | — |
case-11 | pass→pass | 8,824 | 7,010 | -21% | 1 | 1 | 0% | 1,401 | 3,793 | +171% | 0 | 0 | — |
case-15 | fail→fail | 8,188 | 6,008 | -27% | 1 | 1 | 0% | 1,491 | 4,004 | +169% | 0 | 0 | — |
case-16 | pass→fail | 8,335 | 4,775 | -43% | 1 | 1 | 0% | 1,260 | 3,475 | +176% | 0 | 0 | — |
case-17 | pass→pass | 9,514 | 8,391 | -12% | 1 | 1 | 0% | 1,475 | 3,966 | +169% | 0 | 0 | — |
case-18 | pass→pass | 6,683 | 3,941 | -41% | 1 | 1 | 0% | 980 | 3,313 | +238% | 0 | 0 | — |
case-19 | pass→pass | 9,622 | 4,112 | -57% | 1 | 1 | 0% | 1,529 | 3,351 | +119% | 0 | 0 | — |
case-22 | fail→fail | 4,640 | 11,554 | +149% | 1 | 1 | 0% | 181 | 4,661 | +2475% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +23 percentage points is the difference between those two pass rates over the 19 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.