Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use this when the user wants to deeply understand something through guided questioning. Trigger phrases include: "quiz me", "help me understand", "Socratic", "teach me", "walk me through with questions", "test my understanding", or when the user asks for an explanation and would benefit more from guided discovery than a direct answer.
.claude/skills/pchalasani-socratic-quiz/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 28% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -23% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 127% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 15% | 0% |
Guide the user to deep understanding through graduated, adaptive questioning rather than direct explanation. The user learns by thinking through the answers themselves.
understand better (if not already stated).
foundational question — not too easy, not too hard.
response before continuing.
to abstract or nuanced ones.
reason about, or connect to things they already know.
specific behavior, output, or structure they would encounter — but do NOT show them the answer directly.
"why do you think..." style questions.
move to the next, harder question.
stone to the next concept.
gap in their reasoning.
challenges their answer, and ask them to reconsider.
concept, give a small hint (not the answer) and ask again.
incorrect part.
separately, ask a question that requires combining them.
multiple concepts together.
The user should be doing most of the thinking and talking, not you.
give a brief 2-3 sentence summary of what they demonstrated understanding of and what areas might benefit from further exploration.
understanding, not evaluation.
explicitly asks to stop the quiz and just be told.
ask rather than assume.
"That's a really interesting thought!" — just move the conversation forward.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 7,500 | 2,199 | -71% | 1 | 1 | 0% | 1,236 | 1,044 | -16% | 0 | 0 | — |
case-02 | fail→pass | 5,839 | 3,407 | -42% | 1 | 1 | 0% | 990 | 1,264 | +28% | 0 | 0 | — |
case-03 | pass→pass | 4,902 | 4,126 | -16% | 1 | 1 | 0% | 818 | 1,428 | +75% | 0 | 0 | — |
case-04 | pass→pass | 5,488 | 3,526 | -36% | 1 | 1 | 0% | 1,052 | 1,388 | +32% | 0 | 0 | — |
case-05 | fail→fail | 4,718 | 3,283 | -30% | 1 | 1 | 0% | 903 | 1,323 | +47% | 0 | 0 | — |
case-06 | pass→pass | 4,728 | 3,544 | -25% | 1 | 1 | 0% | 928 | 1,302 | +40% | 0 | 0 | — |
case-07 | fail→pass | 9,799 | 4,206 | -57% | 1 | 1 | 0% | 1,781 | 1,372 | -23% | 0 | 0 | — |
case-08 | fail→fail | 2,937 | 3,725 | +27% | 1 | 1 | 0% | 439 | 1,334 | +204% | 0 | 0 | — |
case-09 | pass→pass | 4,642 | 3,629 | -22% | 1 | 1 | 0% | 843 | 1,342 | +59% | 0 | 0 | — |
case-10 | pass→pass | 8,345 | 5,572 | -33% | 1 | 1 | 0% | 1,386 | 1,577 | +14% | 0 | 0 | — |
case-11 | pass→pass | 8,180 | 4,978 | -39% | 1 | 1 | 0% | 1,445 | 1,534 | +6% | 0 | 0 | — |
case-12 | pass→pass | 7,621 | 6,331 | -17% | 1 | 1 | 0% | 1,443 | 1,892 | +31% | 0 | 0 | — |
case-13 | pass→pass | 2,361 | 2,305 | -2% | 1 | 1 | 0% | 407 | 1,186 | +191% | 0 | 0 | — |
case-14 | fail→pass | 3,145 | 2,827 | -10% | 1 | 1 | 0% | 531 | 1,207 | +127% | 0 | 0 | — |
case-15 | fail→pass | 6,226 | 3,325 | -47% | 1 | 1 | 0% | 1,108 | 1,277 | +15% | 0 | 0 | — |
case-16 | pass→pass | 4,018 | 2,590 | -36% | 1 | 1 | 0% | 676 | 1,159 | +71% | 0 | 0 | — |
case-17 | fail→fail | 7,993 | 4,846 | -39% | 1 | 1 | 0% | 1,447 | 1,511 | +4% | 0 | 0 | — |
case-18 | pass→pass | 5,483 | 1,196 | -78% | 1 | 1 | 0% | 1,022 | 906 | -11% | 0 | 0 | — |
case-19 | pass→pass | 4,645 | 2,324 | -50% | 1 | 1 | 0% | 806 | 1,035 | +28% | 0 | 0 | — |
case-20 | fail→pass | 7,106 | 2,937 | -59% | 1 | 1 | 0% | 1,431 | 1,236 | -14% | 0 | 0 | — |
case-21 | fail→pass | 3,933 | 4,452 | +13% | 1 | 1 | 0% | 810 | 1,509 | +86% | 0 | 0 | — |
case-22 | pass→pass | 7,054 | 4,694 | -33% | 1 | 1 | 0% | 1,216 | 1,410 | +16% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +32 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.