Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Evidence-based learning science principles for educational research and practice
.claude/skills/brycewang-stanford-learning-science-guide/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 92% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 20% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -63% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -33% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 30% | 0% |
A comprehensive skill for applying evidence-based learning science principles to educational research, instructional design, and teaching practice. Grounded in cognitive psychology and educational neuroscience.
Working memory has limited capacity. Effective instruction manages three types of cognitive load:
| Load Type | Definition | Design Strategy | |-----------|-----------|-----------------| | Intrinsic | Complexity inherent to the material | Sequence from simple to complex; chunk information | | Extraneous | Load from poor instructional design | Eliminate redundancy; use spatial contiguity | | Germane | Load from schema construction | Use worked examples; encourage self-explanation |
python# Estimate cognitive load using element interactivity def estimate_intrinsic_load(elements: list, interactions: list) -> str: """ elements: list of knowledge components interactions: list of (element_i, element_j) tuples that must be processed simultaneously """ interactivity = len(interactions) / max(len(elements), 1) if interactivity < 0.3: return "low intrinsic load - suitable for independent study" elif interactivity < 0.7: return "moderate intrinsic load - scaffold with worked examples" else: return "high intrinsic load - use fading strategy and segmenting" # Example: teaching statistical regression elements = ['variable', 'coefficient', 'intercept', 'residual', 'R-squared'] interactions = [('coefficient', 'variable'), ('intercept', 'residual'), ('coefficient', 'R-squared'), ('residual', 'R-squared')] print(estimate_intrinsic_load(elements, interactions))
Constructivist approaches emphasize that learners build knowledge through experience. Key active learning strategies with measured effect sizes (Freeman et al., 2014, PNAS):
Testing is not just assessment -- it is a powerful learning tool (Roediger & Karpicke, 2006). Implement the testing effect:
Study Session Structure:
1. Initial encoding (read/watch) - 15 min
2. Free recall (close materials, write) - 10 min
3. Check accuracy and fill gaps - 5 min
4. Spaced retrieval after 1 day - 10 min
5. Spaced retrieval after 7 days - 10 min
6. Spaced retrieval after 30 days - 10 minImplement optimal review scheduling:
pythondef next_review_interval(repetition: int, ease_factor: float = 2.5, quality: int = 4) -> float: """ SM-2 inspired algorithm. repetition: number of successful reviews ease_factor: item difficulty (>= 1.3) quality: response quality 0-5 """ if quality < 3: return 1 # reset to 1 day if repetition == 0: return 1 elif repetition == 1: return 6 else: interval = 6 * (ease_factor ** (repetition - 1)) # Adjust ease factor new_ef = ease_factor + (0.1 - (5 - quality) * (0.08 + (5 - quality) * 0.02)) return round(interval, 1) # Schedule for a moderately difficult concept for rep in range(6): days = next_review_interval(rep) print(f"Review {rep + 1}: after {days} days")
Research shows interleaved practice (mixing problem types) outperforms blocked practice for long-term retention (Rohrer & Taylor, 2007):
Map learning objectives to assessment items across cognitive levels:
yamlremember: verbs: [define, list, recall, identify] assessment: "Multiple choice, matching" understand: verbs: [explain, summarize, compare, classify] assessment: "Short answer, concept maps" apply: verbs: [solve, demonstrate, use, implement] assessment: "Problem sets, simulations" analyze: verbs: [differentiate, organize, attribute, deconstruct] assessment: "Case studies, data interpretation" evaluate: verbs: [judge, critique, justify, appraise] assessment: "Peer review, rubric-based essays" create: verbs: [design, construct, produce, formulate] assessment: "Research projects, portfolios"
After administering assessments, compute item difficulty (p-value) and discrimination index to validate question quality. Target p-values between 0.30 and 0.70 and discrimination indices above 0.30 for optimal measurement.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 10,441 | 14,654 | +40% | 1 | 1 | 0% | 1,572 | 2,522 | +60% | 0 | 0 | — |
case-02 | fail→pass | 16,150 | 23,241 | +44% | 1 | 1 | 0% | 2,636 | 5,071 | +92% | 0 | 0 | — |
case-03 | pass→pass | 2,276 | 1,853 | -19% | 1 | 1 | 0% | 408 | 1,708 | +319% | 0 | 0 | — |
case-04 | pass→pass | 8,626 | 9,626 | +12% | 1 | 1 | 0% | 1,423 | 2,893 | +103% | 0 | 0 | — |
case-05 | pass→pass | 15,372 | 16,666 | +8% | 1 | 1 | 0% | 2,433 | 3,817 | +57% | 0 | 0 | — |
case-06 | pass→pass | 7,721 | 10,451 | +35% | 1 | 1 | 0% | 1,225 | 2,866 | +134% | 0 | 0 | — |
case-07 | pass→pass | 18,478 | 18,557 | +0% | 1 | 1 | 0% | 2,937 | 4,192 | +43% | 0 | 0 | — |
case-08 | fail→pass | 16,421 | 9,315 | -43% | 1 | 1 | 0% | 2,578 | 3,087 | +20% | 0 | 0 | — |
case-09 | fail→pass | 29,619 | 2,903 | -90% | 1 | 1 | 0% | 5,046 | 1,864 | -63% | 0 | 0 | — |
case-10 | fail→pass | 19,447 | 6,002 | -69% | 1 | 1 | 0% | 3,332 | 2,225 | -33% | 0 | 0 | — |
case-11 | fail→pass | 14,004 | 5,606 | -60% | 1 | 1 | 0% | 1,892 | 2,456 | +30% | 0 | 0 | — |
case-12 | pass→pass | 6,777 | 5,431 | -20% | 1 | 1 | 0% | 1,336 | 2,536 | +90% | 0 | 0 | — |
case-13 | pass→pass | 12,610 | 8,725 | -31% | 1 | 1 | 0% | 2,670 | 3,292 | +23% | 0 | 0 | — |
case-14 | pass→pass | 5,953 | 3,747 | -37% | 1 | 1 | 0% | 990 | 1,983 | +100% | 0 | 0 | — |
case-15 | pass→pass | 13,490 | 8,861 | -34% | 1 | 1 | 0% | 2,030 | 2,859 | +41% | 0 | 0 | — |
case-16 | pass→pass | 5,727 | 4,479 | -22% | 1 | 1 | 0% | 903 | 2,049 | +127% | 0 | 0 | — |
case-17 | pass→pass | 5,667 | 3,947 | -30% | 1 | 1 | 0% | 875 | 1,996 | +128% | 0 | 0 | — |
case-18 | pass→pass | 8,495 | 6,120 | -28% | 1 | 1 | 0% | 1,405 | 2,388 | +70% | 0 | 0 | — |
case-19 | fail→pass | 14,930 | 1,943 | -87% | 1 | 1 | 0% | 2,534 | 1,722 | -32% | 0 | 0 | — |
case-20 | pass→pass | 13,321 | 3,389 | -75% | 1 | 1 | 0% | 2,259 | 1,931 | -15% | 0 | 0 | — |
case-21 | pass→pass | 11,518 | 3,935 | -66% | 1 | 1 | 0% | 1,692 | 2,010 | +19% | 0 | 0 | — |
case-22 | pass→pass | 10,714 | 6,252 | -42% | 1 | 1 | 0% | 1,586 | 2,192 | +38% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +27 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.