Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Scaffold a graded problem set with sections, problems, worked solutions, and short "why this matters" explainers across analytical, empirical, and coding types. Use when user says "make a problem set on X", "scaffold exercises for this lecture", "create practice problems", "generate homework with a solution key", "build a graded assignment on topic Y". Emits a clean student set plus a separate solution key — NOT for grading submissions or auto-checking student answers.
.claude/skills/pedrohcgs-scaffold-exercises/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -13% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 126% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 19% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 26% | 0% |
/scaffold-exercises — Problem Set ScaffolderGenerate a graded problem set as two files: a clean student set (problems only) and a solution key (worked solutions + a one-line explainer per problem). Pattern imported from mattpocock/skills, adapted for economics teaching — the primary lens is graded coursework that mixes derivation, estimation, and code.
Input: $ARGUMENTS — a topic (e.g., "instrumental variables", "consumer theory", "quantile regression") and optional flags. See Flags.
Do not use this to grade submissions, auto-check answers, or build a timed exam — it scaffolds practice/graded material, not assessment infrastructure.
| Type | What the student does | Solution artifact | | --- | --- | --- | | analytical | Derive / prove / characterize (theory: optimization, identification, comparative statics) | Step-by-step derivation with the key lemma named | | empirical | Estimate + interpret on a provided or simulated dataset | Expected estimate, sign/magnitude reasoning, common-mistake note | | coding | Implement an estimator or simulation in R or Stata | Runnable reference snippet + expected output shape |
If no dataset is supplied for an empirical problem, generate a small simulated one with a fixed seed (YYYYMMDD) so the answer key is deterministic and reproducible.
Read any source material the user points at (lecture .tex/.qmd, a paper, a dataset header) and produce a Pre-Flight Report before generating problems:
markdown## Pre-Flight Report — Problem Set **Topic:** [topic] **Source(s) read:** [lecture/paper/dataset — one-line takeaway each] **Difficulty:** intro | core | advanced **Counts by type:** analytical=N, empirical=N, coding=N (total = `--count`) **Dataset:** [provided path | simulated with seed YYYYMMDD | none] **Learning objectives:** [2-4 bullets the set should exercise]
Resolve every flag here (interactive choices are gathered before generation, not mid-run). If the topic is too vague to write objectives, ask one clarifying question and stop. Otherwise proceed.
For each problem, write a number, a section heading, the prompt, and any data/notation it needs. Conventions:
create-lecture's pedagogy.For every problem, write:
Bash + R/Stata are available, execute the snippet and paste real output.Emit two files (paths configurable; default under the working directory):
exercises/<topic-slug>_problems.md — the student set: sections, problems, any data, NO answers.exercises/<topic-slug>_solutions.md — the solution key: each problem restated, its worked solution, and its explainer.The split is load-bearing: never leak a solution into the student file. With --no-solutions, write only the student set and stop.
Student set:
markdown# Problem Set: [Topic] (Difficulty: core) ## Section 1 — Analytical **1.** [Motivation sentence.] [Prompt.] ## Section 2 — Empirical **2.** Using `data/<file>` (vars: ...), [estimate + interpret prompt]. ## Section 3 — Coding (R) **3.** [Implement-X prompt.]
Solution key mirrors the numbering, adding ### Solution and > Why this matters: blocks per problem. Close your chat reply with a one-line manifest: files written, problem count by type, and whether code solutions were executed or only drafted.
--difficulty — intro | core | advanced (default core); calibrates step depth as in Phase 1.--count — total number of problems (default 6); split across types per the Pre-Flight counts.--types — comma-separated subset of analytical,empirical,coding (default all three).--dataset — path to a real dataset for empirical problems; omit to simulate one with a seeded DGP.--no-solutions — write only the student set; skip the solution key (Phase 2/3 key file)..claude/skills/create-lecture/SKILL.md — build the lecture these exercises practice; shares notation-reuse + motivation-first conventions..claude/skills/data-analysis/SKILL.md — for empirical problems whose reference solution needs a full R estimation pipeline..claude/skills/simulation-study/SKILL.md — when a problem demonstrates an estimator's finite-sample behavior; reuse its seeded-DGP discipline..claude/skills/lit-review/SKILL.md — source advanced problems from current papers on the topic..claude/skills/interview-me/SKILL.md — turn a fuzzy "I want a set on…" into concrete learning objectives first.templates/skill-template.md — house style for authoring/extending this skill./deploy).| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 71,875 | 77,584 | +8% | 1 | 1 | 0% | 8,293 | 10,028 | +21% | 0 | 0 | — |
case-02 | fail→fail | 77,061 | 43,234 | -44% | 1 | 1 | 0% | 8,212 | 10,013 | +22% | 0 | 0 | — |
case-03 | fail→fail | 40,808 | 41,021 | +1% | 1 | 1 | 0% | 8,268 | 10,003 | +21% | 0 | 0 | — |
case-04 | fail→pass | 11,070 | 5,657 | -49% | 1 | 1 | 0% | 1,974 | 2,568 | +30% | 0 | 0 | — |
case-05 | fail→pass | 13,911 | 3,428 | -75% | 1 | 1 | 0% | 2,613 | 2,281 | -13% | 0 | 0 | — |
case-06 | fail→fail | 16,870 | 12,131 | -28% | 1 | 1 | 0% | 2,690 | 3,643 | +35% | 0 | 0 | — |
case-07 | fail→pass | 22,555 | 43,236 | +92% | 1 | 1 | 0% | 4,409 | 9,975 | +126% | 0 | 0 | — |
case-08 | fail→pass | 21,191 | 5,971 | -72% | 1 | 1 | 0% | 2,321 | 2,770 | +19% | 0 | 0 | — |
case-09 | fail→pass | 14,249 | 9,084 | -36% | 1 | 1 | 0% | 2,653 | 3,347 | +26% | 0 | 0 | — |
case-10 | pass→pass | 16,607 | 27,075 | +63% | 1 | 1 | 0% | 3,262 | 6,877 | +111% | 0 | 0 | — |
case-11 | fail→pass | 11,223 | 21,454 | +91% | 1 | 1 | 0% | 1,943 | 6,118 | +215% | 0 | 0 | — |
case-12 | pass→pass | 33,519 | 40,695 | +21% | 1 | 1 | 0% | 6,482 | 9,951 | +54% | 0 | 0 | — |
case-13 | pass→pass | 3,011 | 6,313 | +110% | 1 | 1 | 0% | 498 | 2,756 | +453% | 0 | 0 | — |
case-14 | pass→pass | 23,554 | 42,936 | +82% | 1 | 1 | 0% | 3,941 | 9,951 | +152% | 0 | 0 | — |
case-15 | fail→pass | 12,226 | 12,398 | +1% | 1 | 1 | 0% | 1,885 | 3,770 | +100% | 0 | 0 | — |
case-16 | pass→pass | 7,816 | 6,090 | -22% | 1 | 1 | 0% | 1,333 | 2,769 | +108% | 0 | 0 | — |
case-17 | pass→pass | 13,199 | 12,313 | -7% | 1 | 1 | 0% | 2,595 | 3,843 | +48% | 0 | 0 | — |
case-18 | fail→pass | 14,832 | 7,986 | -46% | 1 | 1 | 0% | 2,695 | 3,068 | +14% | 0 | 0 | — |
case-19 | fail→fail | 15,019 | 31,350 | +109% | 1 | 1 | 0% | 2,478 | 2,708 | +9% | 0 | 0 | — |
case-20 | fail→fail | 16,436 | 3,190 | -81% | 1 | 1 | 0% | 2,577 | 2,216 | -14% | 0 | 0 | — |
case-21 | fail→fail | 8,815 | 3,148 | -64% | 1 | 1 | 0% | 1,407 | 2,298 | +63% | 0 | 0 | — |
case-22 | fail→pass | 5,754 | 9,739 | +69% | 1 | 1 | 0% | 763 | 3,417 | +348% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +41 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/13/2026 | +36% |
Other measured skills in the registry, with their headline benchmark lift.