Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at study design and again when laying out the results section, to check every slot of a life-science study is filled. Covers the skeleton a paper of this kind carries, and what to put in the slots this run cannot compute rather than leaving them out.
.claude/skills/tangxiangru-life-full-study-skeleton-including-the-wet-lab-half/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 20% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 309% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 569% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 230% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 117% | 0% |
A life-science results section is a fixed skeleton and each slot is judged on its own. Budget one figure per slot.
A pipeline schematic - source data, design or selection, experiment, model, next iteration - drawn before any results figure.
The distribution over biological units: violin or strip of the per-cell, per-donor, per-position value with the unit on the axis, never a mean alone. Name which units behave differently and why; between-unit heterogeneity is itself the finding.
The headline metric stratified along every axis the biology distinguishes - platform or chemistry, cell line or tissue, species or strain, substrate or matrix, gene class, coverage depth, donor - with the pooled figure as one extra row, and say which stratum breaks.
The mechanism in the field's physical vocabulary: which residues, monomers, bases or positions, through which interaction (hydrophobic, electrostatic, steric, stacking), and the measurement that supports it.
Two rules outsiders break. When the supplied material is small, mislabelled, degenerate or a curated extract, still produce the standard analysis on it exactly as given, at the scale given, clearly labelled - then add your corrected analysis and the provenance audit alongside. An audit or an upstream re-scoping that replaces the requested figure loses the result. And the non-computational half - upstream corpus or cohort curation, synthesis and sample-preparation conditions, post-hoc characterisation (mechanical properties, generality across matrices, stability over months) - is compiled from the literature into a labelled comparison table with its provenance stated, never dropped because you could not recompute it.
Attacks the three mechanisms that produce Life's 55% absent rate. (1) Non-computable criteria silently dropped: in one task the five non-computational criteria averaged 7.6 against 43 for the computable ones - this skill makes the literature-compiled table the expected deliverable. (2) The upstream re-scoping / critique-instead-of-deliverable pivot that cost two criteria scoring 3 and 8 on figures that were direct plots of shipped files. (3) Scale mismatch answered with a limitations sentence: 'at the scale given' plus the per-unit distribution slot forces the violin and the stratified table that scored 0. The schematic and mechanism slots cover 4 more criteria the reports never opened a section for.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 51,417 | 55,977 | +9% | 1 | 1 | 0% | 8,358 | 10,065 | +20% | 0 | 0 | — |
case-02 | fail→fail | 63,309 | 51,163 | -19% | 1 | 1 | 0% | 8,275 | 8,789 | +6% | 0 | 0 | — |
case-03 | fail→fail | 48,526 | 49,334 | +2% | 1 | 1 | 0% | 8,272 | 8,786 | +6% | 0 | 0 | — |
case-04 | fail→pass | 13,202 | 73,216 | +455% | 1 | 1 | 0% | 2,085 | 8,527 | +309% | 0 | 0 | — |
case-05 | fail→pass | 57,402 | 50,001 | -13% | 1 | 1 | 0% | 1,230 | 8,234 | +569% | 0 | 0 | — |
case-06 | pass→pass | 17,972 | 50,798 | +183% | 1 | 1 | 0% | 2,990 | 8,744 | +192% | 0 | 0 | — |
case-07 | fail→pass | 14,782 | 43,132 | +192% | 1 | 1 | 0% | 2,199 | 7,248 | +230% | 0 | 0 | — |
case-08 | fail→pass | 15,871 | 30,024 | +89% | 1 | 1 | 0% | 2,226 | 4,829 | +117% | 0 | 0 | — |
case-09 | fail→pass | 12,385 | 36,643 | +196% | 1 | 1 | 0% | 1,737 | 6,404 | +269% | 0 | 0 | — |
case-10 | fail→pass | 14,944 | 42,918 | +187% | 1 | 1 | 0% | 2,195 | 7,036 | +221% | 0 | 0 | — |
case-11 | fail→pass | 19,250 | 53,539 | +178% | 1 | 1 | 0% | 2,982 | 7,222 | +142% | 0 | 0 | — |
case-12 | fail→pass | 13,845 | 42,979 | +210% | 1 | 1 | 0% | 1,958 | 7,101 | +263% | 0 | 0 | — |
case-13 | fail→pass | 13,293 | 51,238 | +285% | 1 | 1 | 0% | 1,374 | 8,745 | +536% | 0 | 0 | — |
case-14 | fail→pass | 20,613 | 35,713 | +73% | 1 | 1 | 0% | 3,062 | 6,151 | +101% | 0 | 0 | — |
case-15 | fail→pass | 15,614 | 36,640 | +135% | 1 | 1 | 0% | 2,155 | 6,082 | +182% | 0 | 0 | — |
case-16 | fail→fail | 13,870 | 50,789 | +266% | 1 | 1 | 0% | 2,060 | 8,745 | +325% | 0 | 0 | — |
case-17 | fail→pass | 5,324 | 41,229 | +674% | 1 | 1 | 0% | 629 | 7,166 | +1039% | 0 | 0 | — |
case-18 | fail→fail | 4,798 | 36,898 | +669% | 1 | 1 | 0% | 545 | 6,065 | +1013% | 0 | 0 | — |
case-19 | fail→pass | 19,244 | 41,432 | +115% | 1 | 1 | 0% | 2,721 | 7,335 | +170% | 0 | 0 | — |
case-20 | fail→pass | 10,160 | 35,353 | +248% | 1 | 1 | 0% | 1,428 | 6,086 | +326% | 0 | 0 | — |
case-21 | fail→fail | 23,571 | 49,416 | +110% | 1 | 1 | 0% | 4,191 | 8,017 | +91% | 0 | 0 | — |
case-22 | fail→pass | 62,790 | 36,771 | -41% | 1 | 1 | 0% | 4,124 | 6,796 | +65% | 0 | 0 | — |
case-23 | pass→fail | 20,899 | 48,288 | +131% | 1 | 1 | 0% | 3,134 | 8,728 | +178% | 0 | 0 | — |
case-24 | pass→fail | 17,430 | 72,049 | +313% | 1 | 1 | 0% | 3,292 | 7,245 | +120% | 0 | 0 | — |
case-25 | pass→fail | 16,892 | 43,033 | +155% | 1 | 1 | 0% | 2,120 | 7,106 | +235% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 24 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +52 percentage points is the difference between those two pass rates over the 24 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.