Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at analysis and figure planning when the brief names the units its data is grouped into — patients, cells, classes, labs, behaviours — and you are about to report one pooled number over all of them. Covers why the pooled number hides the result, which strata a study of this kind is expected to report, and when an aggregate is the right answer after all.
.claude/skills/tangxiangru-the-unit-of-analysis/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 140% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -7% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 11% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 175% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -20% | 0% |
A single pooled number is the easiest result to compute and often the least informative one. If the data distinguishes patients, regions, epochs, conditions, lead times, molecules or seeds, then the analysis is expected at that level — and a pooled mean is read as the stratified analysis not having been done.
The failure is specific and common: the study has seven subjects, the report shows one distribution over all of them, and the question the study existed to answer — do subjects differ, and how — is unanswerable from the figure.
site, condition or fold a row belongs to. It is almost always there.
distribution with the units as points, or a table with a row each.
That is usually the most interesting sentence in the report.
When the claim is genuinely about the population, and you have shown the per-unit spread somewhere so the reader can judge whether pooling was fair. A pooled number whose spread is never shown asks to be trusted rather than read.
Look at your main figure. Can a reader tell how many units contributed and whether they agreed? If not, you have shown a summary and called it a result.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 8,614 | 16,962 | +97% | 1 | 1 | 0% | 1,150 | 2,759 | +140% | 0 | 0 | — |
case-02 | fail→pass | 37,747 | 15,346 | -59% | 1 | 1 | 0% | 2,584 | 2,415 | -7% | 0 | 0 | — |
case-03 | pass→pass | 13,638 | 24,041 | +76% | 1 | 1 | 0% | 2,226 | 2,449 | +10% | 0 | 0 | — |
case-04 | pass→pass | 22,263 | 24,540 | +10% | 1 | 1 | 0% | 3,594 | 3,031 | -16% | 0 | 0 | — |
case-05 | fail→pass | 19,379 | 19,525 | +1% | 1 | 1 | 0% | 3,045 | 3,384 | +11% | 0 | 0 | — |
case-06 | pass→pass | 16,333 | 10,096 | -38% | 1 | 1 | 0% | 2,509 | 2,035 | -19% | 0 | 0 | — |
case-07 | fail→pass | 5,893 | 13,344 | +126% | 1 | 1 | 0% | 852 | 2,339 | +175% | 0 | 0 | — |
case-08 | pass→pass | 26,758 | 18,758 | -30% | 1 | 1 | 0% | 3,031 | 3,223 | +6% | 0 | 0 | — |
case-09 | fail→pass | 32,263 | 12,315 | -62% | 1 | 1 | 0% | 2,738 | 2,204 | -20% | 0 | 0 | — |
case-10 | pass→pass | 27,699 | 21,218 | -23% | 1 | 1 | 0% | 3,557 | 3,728 | +5% | 0 | 0 | — |
case-11 | pass→pass | 183,135 | 12,427 | -93% | 1 | 1 | 0% | 2,981 | 2,302 | -23% | 0 | 0 | — |
case-12 | fail→pass | 14,709 | 10,600 | -28% | 1 | 1 | 0% | 2,260 | 1,940 | -14% | 0 | 0 | — |
case-13 | fail→pass | 32,003 | 20,893 | -35% | 1 | 1 | 0% | 2,863 | 2,574 | -10% | 0 | 0 | — |
case-14 | pass→pass | 17,556 | 22,088 | +26% | 1 | 1 | 0% | 2,871 | 2,891 | +1% | 0 | 0 | — |
case-15 | pass→pass | 19,876 | 34,236 | +72% | 1 | 1 | 0% | 3,150 | 2,474 | -21% | 0 | 0 | — |
case-16 | pass→pass | 20,374 | 15,988 | -22% | 1 | 1 | 0% | 3,048 | 2,460 | -19% | 0 | 0 | — |
case-17 | fail→pass | 15,932 | 14,174 | -11% | 1 | 1 | 0% | 2,690 | 2,622 | -3% | 0 | 0 | — |
case-18 | fail→pass | 34,512 | 147,812 | +328% | 1 | 1 | 0% | 2,437 | 3,195 | +31% | 0 | 0 | — |
case-19 | fail→pass | 25,747 | 50,608 | +97% | 1 | 1 | 0% | 1,937 | 3,605 | +86% | 0 | 0 | — |
case-20 | pass→pass | 18,467 | 15,804 | -14% | 1 | 1 | 0% | 2,979 | 3,177 | +7% | 0 | 0 | — |
case-21 | pass→pass | 68,107 | 12,243 | -82% | 1 | 1 | 0% | 2,252 | 2,223 | -1% | 0 | 0 | — |
case-22 | pass→pass | 18,103 | 19,935 | +10% | 1 | 1 | 0% | 3,258 | 3,920 | +20% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +45 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.