Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at analysis when a detection or classification result is about to be reported as one accuracy over a pooled population. Covers per-group and per-class precision, recall and confusion matrices at a stated threshold, and sweeping the degradations the recording modality actually suffers.
.claude/skills/tangxiangru-neuroscience-stratify-and-report-detection-metrics/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 8% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 298% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 109% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 41% | 0% |
Neuroscience data arrives with grouping factors, and the field reports per group. Before analysing, list every categorical in the design or the file that is not the label - experimental condition, subject or session, region, cell type, acquisition site, artifact class - and make each one a reporting axis. Every headline number is reported per level, with the pooled value as one additional row. A column present in the data but absent from your tables reads as an unreported factor.
Events of interest are usually rare, so use the detection metric family rather than accuracy or AUROC: precision, recall and F1 at a stated operating threshold, a precision-recall curve with average precision, and a confusion matrix - per class and per stratum. AUROC is insensitive to the prevalence regime the science operates in; report it only alongside these.
Robustness means the corruptions the instrument itself produces - motion, drift, channel or electrode loss, line noise, low SNR, downsampling, session-to-session shift - swept one at a time, each as a degradation curve of the same metric, with the comparison methods on the same axes so relative decay rates are visible. Generic added Gaussian noise does not test this.
Open with provenance: source dataset, acquisition modality, physical extent (volume, duration, subjects, cells), how ground-truth labels were obtained, and how it compares in scale and diversity to the field's other benchmark datasets. If your input is a reduced or simulated stand-in, say what it stands for and instantiate the analyses on it anyway.
Four of four Neuroscience tasks require per-stratum reporting and it is the single most-missed axis: agents stratified by their own methodological axis (evaluation protocol, learner family) rather than the domain column shipped in the data, and where per-stratum precision/recall/F1 was demanded they reported pooled AUROC/AP. One run computed per-degradation-stratum AUROC into a CSV and put no per-stratum table in the report at all. The provenance clause plus 'instantiate the analyses anyway' attacks the dominant proxy-data pivot: all 8 absent criteria sit in the two tasks whose input is a reduced stand-in, where the agent replaced the study with a critique of the stand-in.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 48,953 | 43,600 | -11% | 1 | 1 | 0% | 8,311 | 8,784 | +6% | 0 | 0 | — |
case-02 | fail→fail | 69,463 | 41,953 | -40% | 1 | 1 | 0% | 8,587 | 8,771 | +2% | 0 | 0 | — |
case-03 | fail→pass | 44,241 | 41,588 | -6% | 1 | 1 | 0% | 8,105 | 8,765 | +8% | 0 | 0 | — |
case-04 | fail→fail | 27,293 | 59,197 | +117% | 1 | 1 | 0% | 2,949 | 8,936 | +203% | 0 | 0 | — |
case-05 | fail→pass | 17,495 | 47,384 | +171% | 1 | 1 | 0% | 2,018 | 8,031 | +298% | 0 | 0 | — |
case-06 | fail→fail | 30,081 | 44,493 | +48% | 1 | 1 | 0% | 3,986 | 8,741 | +119% | 0 | 0 | — |
case-07 | fail→fail | 13,908 | 25,527 | +84% | 1 | 1 | 0% | 2,022 | 3,266 | +62% | 0 | 0 | — |
case-08 | fail→fail | 30,562 | 32,057 | +5% | 1 | 1 | 0% | 1,261 | 6,702 | +431% | 0 | 0 | — |
case-09 | fail→pass | 17,838 | 46,792 | +162% | 1 | 1 | 0% | 2,842 | 5,948 | +109% | 0 | 0 | — |
case-10 | fail→fail | 19,895 | 36,349 | +83% | 1 | 1 | 0% | 3,063 | 6,507 | +112% | 0 | 0 | — |
case-11 | fail→pass | 22,821 | 38,071 | +67% | 1 | 1 | 0% | 3,598 | 5,071 | +41% | 0 | 0 | — |
case-12 | fail→pass | 22,769 | 23,349 | +3% | 1 | 1 | 0% | 3,739 | 4,300 | +15% | 0 | 0 | — |
case-13 | fail→fail | 29,602 | 13,049 | -56% | 1 | 1 | 0% | 2,234 | 2,530 | +13% | 0 | 0 | — |
case-14 | fail→pass | 18,931 | 37,890 | +100% | 1 | 1 | 0% | 2,813 | 4,256 | +51% | 0 | 0 | — |
case-15 | fail→fail | 35,506 | 48,084 | +35% | 1 | 1 | 0% | 3,092 | 4,191 | +36% | 0 | 0 | — |
case-16 | fail→pass | 9,418 | 18,776 | +99% | 1 | 1 | 0% | 1,337 | 3,327 | +149% | 0 | 0 | — |
case-17 | fail→pass | 18,318 | 14,591 | -20% | 1 | 1 | 0% | 2,493 | 2,795 | +12% | 0 | 0 | — |
case-18 | fail→pass | 20,583 | 55,872 | +171% | 1 | 1 | 0% | 3,103 | 5,576 | +80% | 0 | 0 | — |
case-19 | fail→pass | 28,381 | 29,193 | +3% | 1 | 1 | 0% | 2,544 | 2,720 | +7% | 0 | 0 | — |
case-20 | pass→fail | 18,651 | 42,526 | +128% | 1 | 1 | 0% | 3,367 | 8,731 | +159% | 0 | 0 | — |
case-21 | pass→fail | 35,588 | 43,472 | +22% | 1 | 1 | 0% | 3,420 | 8,712 | +155% | 0 | 0 | — |
case-22 | pass→fail | 28,553 | 40,671 | +42% | 1 | 1 | 0% | 1,942 | 8,547 | +340% | 0 | 0 | — |
case-23 | pass→fail | 22,629 | 45,964 | +103% | 1 | 1 | 0% | 1,997 | 9,524 | +377% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +30 percentage points is the difference between those two pass rates over the 22 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.