Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at study design, before any experiment of your own is costed, on reproduction and method-evaluation tasks. Covers enumerating the systems, scenarios, stress sweeps and case studies the source names, running each one by name, measuring the preconditions the method declares it needs, and what to do when one of them fails.
.claude/skills/tangxiangru-run-the-conditions-the-source-ran/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -2% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -5% | 0% |
A paper's results are attached to specific things: named systems, named scenarios, a named molecule, a named problem, a named noise axis, and the conditions the method says it needs. A checklist for reproducing that paper is written from those names. A better-designed experiment on different conditions answers a question nobody asked.
A run reproducing a feature-selection method replaced the paper's robustness experiment — degradation under falling signal-to-noise, reduced library size and increased dropout, against two named baselines — with its own structured confounder and batch-geometry design. Better science, in the abstract, and the reviewer wrote: "it does not perform or report any simulations varying SNR, library size, or dropout, nor explicitly compare performance degradation curves versus Laplacian Score or MCFS." That requirement scored 5. A plain agent that simply ran the paper's sweep scored 65 on the same requirement.
The same shape recurs, and it is never laziness — it is always a substitution made for a good local reason:
worth computing, and argued the point instead of computing them on the paper's own molecule: 5 against 70.
the requirement that names it scored 12 against 38.
paper's worked example, IMO 2004 P1: 5 against 18.
method's flat-ground assumption — correlation 0.146, camera height constant within each track — and then never mentioned it again. The word "perspective" occurs 0 times in its report and 4 times in each comparator's.
notes/source_experiments.json: one row per namedsystem, sample, scenario, stress axis, ablation and worked case study, each carrying the baselines it is compared against and the numbers the source reports for it. Add a row for every condition the method declares it needs — a ground plane, i.i.d. residuals, sparsity, a calibrated confidence channel.
the section heading. Your improved design is an extra row, never a replacement. If the budget cannot carry both, cut your variant.
where the mechanism is introduced — not in Limitations. A precondition that fails is one of the strongest things a reproduction can find, and it is worth nothing if the reader meets it as an apology on the last page.
data where it holds, re-run the same comparison there, and report the pair. That turns "the assumption does not hold" into "the assumption does not hold, and here is what it costs" — which is the finding.
you did instead, and what the gap is. Silence on a named experiment reads as an experiment nobody thought of.
Walk source_experiments.json. Every row should map to a heading in the report that uses the source's own name for it. A row whose answer is "we did something better" is the failure above, wearing its best clothes.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 32,080 | 39,894 | +24% | 1 | 1 | 0% | 5,565 | 4,684 | -16% | 0 | 0 | — |
case-02 | fail→pass | 32,519 | 18,766 | -42% | 1 | 1 | 0% | 4,033 | 3,933 | -2% | 0 | 0 | — |
case-03 | fail→pass | 17,313 | 18,430 | +6% | 1 | 1 | 0% | 2,932 | 3,547 | +21% | 0 | 0 | — |
case-04 | pass→pass | 12,710 | 21,284 | +67% | 1 | 1 | 0% | 2,246 | 2,600 | +16% | 0 | 0 | — |
case-05 | fail→pass | 18,971 | 19,521 | +3% | 1 | 1 | 0% | 3,389 | 3,959 | +17% | 0 | 0 | — |
case-06 | pass→pass | 18,814 | 180,435 | +859% | 1 | 1 | 0% | 2,954 | 5,073 | +72% | 0 | 0 | — |
case-07 | fail→pass | 39,754 | 11,214 | -72% | 1 | 1 | 0% | 2,370 | 2,254 | -5% | 0 | 0 | — |
case-08 | fail→pass | 21,112 | 146,981 | +596% | 1 | 1 | 0% | 2,952 | 2,739 | -7% | 0 | 0 | — |
case-09 | fail→pass | 13,566 | 6,892 | -49% | 1 | 1 | 0% | 1,968 | 1,734 | -12% | 0 | 0 | — |
case-10 | fail→pass | 12,520 | 18,617 | +49% | 1 | 1 | 0% | 1,908 | 2,300 | +21% | 0 | 0 | — |
case-11 | fail→pass | 14,553 | 11,663 | -20% | 1 | 1 | 0% | 1,986 | 2,525 | +27% | 0 | 0 | — |
case-12 | fail→pass | 30,719 | 28,010 | -9% | 1 | 1 | 0% | 3,086 | 3,271 | +6% | 0 | 0 | — |
case-13 | fail→pass | 35,109 | 10,496 | -70% | 1 | 1 | 0% | 1,951 | 2,252 | +15% | 0 | 0 | — |
case-14 | fail→pass | 21,456 | 14,123 | -34% | 1 | 1 | 0% | 1,983 | 1,889 | -5% | 0 | 0 | — |
case-15 | pass→pass | 19,744 | 8,259 | -58% | 1 | 1 | 0% | 2,204 | 2,063 | -6% | 0 | 0 | — |
case-16 | fail→pass | 15,865 | 10,021 | -37% | 1 | 1 | 0% | 2,394 | 2,166 | -10% | 0 | 0 | — |
case-17 | fail→pass | 14,803 | 17,069 | +15% | 1 | 1 | 0% | 2,440 | 2,306 | -5% | 0 | 0 | — |
case-18 | fail→pass | 13,674 | 11,874 | -13% | 1 | 1 | 0% | 2,038 | 2,144 | +5% | 0 | 0 | — |
case-19 | fail→pass | 14,529 | 8,606 | -41% | 1 | 1 | 0% | 1,973 | 1,927 | -2% | 0 | 0 | — |
case-20 | fail→pass | 20,755 | 31,032 | +50% | 1 | 1 | 0% | 2,807 | 2,874 | +2% | 0 | 0 | — |
case-21 | fail→fail | 26,895 | 28,364 | +5% | 1 | 1 | 0% | 2,605 | 2,211 | -15% | 0 | 0 | — |
case-22 | fail→pass | 8,288 | 6,216 | -25% | 1 | 1 | 0% | 1,032 | 1,630 | +58% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +82 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.