Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at study design and analysis when a model is about to be compared against one alternative, or a fit reported without a negative control. Covers the two-sided comparator ladder, the control representation panel, and splitting per-unit predictions into the ones a measurement validates and the ones that stay predictions.
.claude/skills/tangxiangru-neuroscience-comparator-ladder-and-per-unit-predictions/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -31% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -22% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -21% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 20% | 0% |
Three deliverables a computational-neuroscience claim must carry; plan them before fitting.
A two-sided comparator ladder. Horizontally, a bank of named alternative methods spanning the families in use (supervised, correlation based, variance based, single-modality). Vertically, variants of your own model that each delete one information source your thesis says is necessary - structure or connectivity, task optimisation, temporal order, one modality - plus a granularity sweep that coarsens the entity taxonomy (fine type, family, broad class) until performance collapses. Same metrics, same split, every rung.
The paired representation panel. Show the low-dimensional projection twice on identical axes and colouring: once under the untreated, shuffled or degraded condition and once under the treated one, quantified with the same gap metric in both. The deliberately poor control panel is a required deliverable, not something the good panel excuses.
Per-unit publication. Give the model's per-unit quantity for the whole population as a ranked table or figure, partitioned into units where an independent measurement exists (report agreement as k of n) and units where the model issues an untested prediction, labelled as novel predictions, with the coverage fraction stated. Pair it with a mechanism established by intervention - property P of upstream unit A sets property Q of downstream unit B, shown by silencing or re-weighting A - closing with a prediction an experimentalist could test.
Each must appear as a labelled figure with its numbers in the panel or caption. A quantity computed into a results file but never plotted counts as not done.
Three tasks each demand the comparator ladder, the paired projection with its deliberately-bad control panel, and the intervention-based mechanism statement; three demand the per-unit table split into validated versus novel predictions with a stated coverage fraction. Bare runs ran textbook single-input ablations instead of the variant grid, never ran a taxonomy-coarsening sweep, and never drew the negative-condition panel. The closing rule targets the largest measured single loss in the discipline: per-feature attributions written to a diagnostics JSON, never plotted and never named in the report, taking 0.60 of one task's weight to zero.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 42,277 | 28,189 | -33% | 1 | 1 | 0% | 6,993 | 4,849 | -31% | 0 | 0 | — |
case-02 | fail→pass | 52,973 | 37,993 | -28% | 1 | 1 | 0% | 8,259 | 6,439 | -22% | 0 | 0 | — |
case-03 | fail→pass | 38,090 | 28,907 | -24% | 1 | 1 | 0% | 6,353 | 5,050 | -21% | 0 | 0 | — |
case-04 | fail→pass | 30,401 | 29,039 | -4% | 1 | 1 | 0% | 3,097 | 4,396 | +42% | 0 | 0 | — |
case-05 | fail→pass | 19,918 | 21,256 | +7% | 1 | 1 | 0% | 2,939 | 3,528 | +20% | 0 | 0 | — |
case-06 | pass→pass | 15,255 | 15,437 | +1% | 1 | 1 | 0% | 2,277 | 2,758 | +21% | 0 | 0 | — |
case-07 | fail→pass | 17,471 | 17,649 | +1% | 1 | 1 | 0% | 2,552 | 3,021 | +18% | 0 | 0 | — |
case-08 | fail→pass | 15,193 | 14,590 | -4% | 1 | 1 | 0% | 1,913 | 2,506 | +31% | 0 | 0 | — |
case-09 | fail→pass | 13,298 | 9,147 | -31% | 1 | 1 | 0% | 1,911 | 1,884 | -1% | 0 | 0 | — |
case-10 | fail→pass | 16,596 | 8,712 | -48% | 1 | 1 | 0% | 2,298 | 1,475 | -36% | 0 | 0 | — |
case-11 | pass→pass | 20,657 | 17,448 | -16% | 1 | 1 | 0% | 2,912 | 2,891 | -1% | 0 | 0 | — |
case-12 | fail→pass | 40,574 | 25,513 | -37% | 1 | 1 | 0% | 3,179 | 4,276 | +35% | 0 | 0 | — |
case-13 | fail→pass | 19,209 | 11,264 | -41% | 1 | 1 | 0% | 2,215 | 2,111 | -5% | 0 | 0 | — |
case-14 | fail→pass | 17,322 | 8,152 | -53% | 1 | 1 | 0% | 2,444 | 1,733 | -29% | 0 | 0 | — |
case-15 | pass→pass | 13,922 | 29,110 | +109% | 1 | 1 | 0% | 2,031 | 4,773 | +135% | 0 | 0 | — |
case-16 | pass→pass | 16,044 | 6,692 | -58% | 1 | 1 | 0% | 2,152 | 1,402 | -35% | 0 | 0 | — |
case-17 | fail→pass | 19,893 | 14,650 | -26% | 1 | 1 | 0% | 2,916 | 2,475 | -15% | 0 | 0 | — |
case-18 | fail→pass | 13,734 | 8,401 | -39% | 1 | 1 | 0% | 1,760 | 1,512 | -14% | 0 | 0 | — |
case-19 | pass→pass | 15,794 | 15,359 | -3% | 1 | 1 | 0% | 2,330 | 2,628 | +13% | 0 | 0 | — |
case-20 | fail→pass | 17,653 | 10,664 | -40% | 1 | 1 | 0% | 2,164 | 1,882 | -13% | 0 | 0 | — |
case-21 | fail→pass | 17,018 | 17,872 | +5% | 1 | 1 | 0% | 2,259 | 3,022 | +34% | 0 | 0 | — |
case-22 | pass→pass | 20,799 | 31,699 | +52% | 1 | 1 | 0% | 3,078 | 5,286 | +72% | 0 | 0 | — |
case-23 | pass→pass | 17,698 | 27,594 | +56% | 1 | 1 | 0% | 2,780 | 4,704 | +69% | 0 | 0 | — |
case-24 | pass→pass | 27,316 | 34,776 | +27% | 1 | 1 | 0% | 4,259 | 6,252 | +47% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted. The headline lift of +67 percentage points is the difference between those two pass rates over the 24 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.