Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Design and diagnose list experiments (item count technique).
.claude/skills/brycewang-stanford-list-experiment/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 57% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 81% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 53% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 116% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 74% | 0% |
Related skills: Use alongside hypothesis-building (state π and a SESOI before design choices), survey-design (mode effects, question ordering, and pre-testing of control items), and methods-reporting (deposit list wording, randomization seed, list package version, and ict.test / ict.hausman.test / ictreg() output).
list R package.list R package (Blair, Chou & Imai), which provides a unified interface for difference-in-means, NLSreg, MLreg, combined estimator, and Bayesian MCMC hierarchical models, along with all standard diagnostic tests.ict.test() in Blair & Imai's (2012) list package.ictreg().ict.hausman.test() in the list package — reject model specification if the Hausman statistic is large and positive, or if it takes a negative value (which itself signals misspecification). If detected, use NLSreg as the primary estimator and consider including a placebo item.list package's simulation tools support this. Rule of thumb: assume effective sample sizes 5–10× below what a direct question study would require.ict.test() and reported?ict.hausman.test(); Blair, Chou & Imai 2019) reported when a multivariate estimator is used?list package cited: Is the list R package (Blair, Chou & Imai) cited as the implementation source?For a worked illustration — a four-item control list for a clientelism / vote-buying sensitive item, with expected prevalences, floor/ceiling tail calculations, a pre-field NFC simulation, and the specific ict.test() / ict.hausman.test() diagnostic calls — see reference/example-clientelism.md.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 33,103 | 31,798 | -4% | 1 | 1 | 0% | 6,222 | 9,767 | +57% | 0 | 0 | — |
case-02 | fail→pass | 29,442 | 31,239 | +6% | 1 | 1 | 0% | 5,351 | 9,669 | +81% | 0 | 0 | — |
case-03 | fail→pass | 30,386 | 29,869 | -2% | 1 | 1 | 0% | 6,204 | 9,473 | +53% | 0 | 0 | — |
case-04 | pass→pass | 24,406 | 31,683 | +30% | 1 | 1 | 0% | 4,856 | 9,848 | +103% | 0 | 0 | — |
case-05 | pass→pass | 18,869 | 26,967 | +43% | 1 | 1 | 0% | 3,800 | 9,001 | +137% | 0 | 0 | — |
case-06 | pass→pass | 18,776 | 22,402 | +19% | 1 | 1 | 0% | 3,808 | 8,078 | +112% | 0 | 0 | — |
case-07 | pass→pass | 14,354 | 11,750 | -18% | 1 | 1 | 0% | 2,422 | 5,670 | +134% | 0 | 0 | — |
case-08 | pass→pass | 12,181 | 9,781 | -20% | 1 | 1 | 0% | 1,962 | 5,439 | +177% | 0 | 0 | — |
case-09 | pass→pass | 11,961 | 10,069 | -16% | 1 | 1 | 0% | 2,023 | 5,375 | +166% | 0 | 0 | — |
case-10 | fail→pass | 10,756 | 5,084 | -53% | 1 | 1 | 0% | 2,140 | 4,625 | +116% | 0 | 0 | — |
case-11 | fail→pass | 18,038 | 8,116 | -55% | 1 | 1 | 0% | 2,986 | 5,182 | +74% | 0 | 0 | — |
case-12 | fail→pass | 31,630 | 7,025 | -78% | 1 | 1 | 0% | 1,756 | 5,016 | +186% | 0 | 0 | — |
case-13 | fail→pass | 13,433 | 6,113 | -54% | 1 | 1 | 0% | 2,235 | 4,720 | +111% | 0 | 0 | — |
case-14 | fail→pass | 21,660 | 10,142 | -53% | 1 | 1 | 0% | 3,776 | 5,525 | +46% | 0 | 0 | — |
case-15 | fail→pass | 7,698 | 4,426 | -43% | 1 | 1 | 0% | 1,319 | 4,433 | +236% | 0 | 0 | — |
case-16 | fail→pass | 30,553 | 10,577 | -65% | 1 | 1 | 0% | 2,449 | 5,690 | +132% | 0 | 0 | — |
case-17 | pass→pass | 24,615 | 13,684 | -44% | 1 | 1 | 0% | 3,819 | 6,131 | +61% | 0 | 0 | — |
case-18 | fail→pass | 18,954 | 6,507 | -66% | 1 | 1 | 0% | 3,269 | 4,718 | +44% | 0 | 0 | — |
case-19 | pass→pass | 7,273 | 7,611 | +5% | 1 | 1 | 0% | 1,344 | 4,998 | +272% | 0 | 0 | — |
case-20 | fail→pass | 17,134 | 7,362 | -57% | 1 | 1 | 0% | 2,719 | 4,960 | +82% | 0 | 0 | — |
case-21 | pass→pass | 9,776 | 7,248 | -26% | 1 | 1 | 0% | 1,628 | 4,850 | +198% | 0 | 0 | — |
case-22 | pass→pass | 11,300 | 8,781 | -22% | 1 | 1 | 0% | 1,921 | 5,217 | +172% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +55 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.