Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at Stage 02 and Stage 03 whenever the task is to reproduce, re-implement or verify a published study and the hypotheses you are drafting are all about something else. Covers how to write the reproduction itself as a falsifiable frozen commitment, why a self-invented question crowds it out, and how to budget between the two.
.claude/skills/tangxiangru-the-reproduction-is-a-hypothesis/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 45% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 62% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 44% | 0% |
A hypothesis has to be able to fail. That test is easy to apply badly to a reproduction, because "we will re-implement the paper" is an engineering goal with no plausible failure and no surprise in either direction — so the drafting stage rejects it, and the run goes looking for a question of its own.
The question it finds is often a good one. It is also, reliably, not the task's question, and once it is frozen it becomes the only contract the run has. Every later stage adjudicates it, the figure plan is drawn around it, and the report is organised by it. The reproduction the task actually asked for survives as a subsection, or as a comparison table, or not at all.
The engineering-goal objection is right about the phrasing and wrong about the content. "We will re-implement X" fails the test. These do not:
our estimate falls within the source's stated uncertainty of its value; refuted if it differs by more than that. This can fail, it fails often, and when it fails you have a finding the field would want.
source reports A beats B; our reproduction holds the budget equal and asks whether the ordering survives. Refutable by construction.
explanation predicts a specific quantity behaves a specific way; measure that quantity.
Each of these is the reproduction, and each names an outcome that would surprise someone. A reproduction that lands on the number is a confirmation with a tolerance; one that misses is a discrepancy with a magnitude. Both are results.
When the task names a reproduction, the reproduction's hypotheses are frozen first and get the larger share of the compute. Your own question is an extension, it is frozen second, and it is budgeted from what is left. This ordering is not modesty — it is what makes your extension interpretable. An improvement measured against a reproduction you did not complete cannot be attributed to the improvement.
Count the hypotheses you are about to freeze. If the task's outputs are three things and you are freezing eighteen propositions, none of which is "the three things come out right", the contract you are signing is not the one you were given. Freezing more hypotheses is not more rigour; it is a longer list of questions the grader did not ask.
The literature survey will sometimes establish that exact numeric agreement with the source is not well posed — the seed is unstated, the data has moved, the hardware differs. That is worth recording. It is not a reason to replace the reproduction with a study of why reproduction is hard. Record the obstacle, set a tolerance that accounts for it, and reproduce against the tolerance. A conclusion reached before the first experiment should widen an error bar, never delete an arm.
Read the task statement's list of outputs, then read the frozen hypothesis set. Every named output should be adjudicated by at least one hypothesis whose decision rule mentions a quantity from the source. If the two lists share nothing, go back: the run is about to spend its whole budget proving something nobody asked.
See also reproduce-then-extend for the shape of the comparison, cover-what-the-task-named for enumerating the outputs, and close-the-gap-to-the-published-number for what to do when the reproduction lands off the published value.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 20,320 | 19,616 | -3% | 1 | 1 | 0% | 3,220 | 3,929 | +22% | 0 | 0 | — |
case-02 | pass→pass | 41,961 | 13,447 | -68% | 1 | 1 | 0% | 2,988 | 2,680 | -10% | 0 | 0 | — |
case-03 | fail→pass | 15,119 | 36,752 | +143% | 1 | 1 | 0% | 2,337 | 3,393 | +45% | 0 | 0 | — |
case-04 | fail→pass | 23,333 | 20,182 | -14% | 1 | 1 | 0% | 3,286 | 3,791 | +15% | 0 | 0 | — |
case-05 | fail→pass | 16,699 | 42,477 | +154% | 1 | 1 | 0% | 2,269 | 3,675 | +62% | 0 | 0 | — |
case-06 | pass→pass | 23,263 | 20,051 | -14% | 1 | 1 | 0% | 3,195 | 3,389 | +6% | 0 | 0 | — |
case-07 | pass→pass | 16,235 | 10,718 | -34% | 1 | 1 | 0% | 2,771 | 2,670 | -4% | 0 | 0 | — |
case-08 | pass→pass | 23,603 | 8,764 | -63% | 1 | 1 | 0% | 2,007 | 2,173 | +8% | 0 | 0 | — |
case-09 | fail→pass | 79,022 | 14,709 | -81% | 1 | 1 | 0% | 2,267 | 3,273 | +44% | 0 | 0 | — |
case-10 | fail→pass | 47,033 | 32,697 | -30% | 1 | 1 | 0% | 3,107 | 2,587 | -17% | 0 | 0 | — |
case-11 | fail→pass | 13,732 | 11,460 | -17% | 1 | 1 | 0% | 2,063 | 2,402 | +16% | 0 | 0 | — |
case-12 | fail→pass | 14,029 | 8,648 | -38% | 1 | 1 | 0% | 1,898 | 1,916 | +1% | 0 | 0 | — |
case-13 | fail→pass | 11,324 | 7,751 | -32% | 1 | 1 | 0% | 1,644 | 1,935 | +18% | 0 | 0 | — |
case-14 | pass→pass | 12,967 | 12,872 | -1% | 1 | 1 | 0% | 1,840 | 2,680 | +46% | 0 | 0 | — |
case-15 | pass→pass | 19,435 | 33,871 | +74% | 1 | 1 | 0% | 2,165 | 2,938 | +36% | 0 | 0 | — |
case-16 | pass→pass | 25,300 | 18,495 | -27% | 1 | 1 | 0% | 2,069 | 2,408 | +16% | 0 | 0 | — |
case-17 | fail→pass | 27,743 | 22,318 | -20% | 1 | 1 | 0% | 2,521 | 2,912 | +16% | 0 | 0 | — |
case-18 | fail→fail | 22,982 | 11,463 | -50% | 1 | 1 | 0% | 2,400 | 2,428 | +1% | 0 | 0 | — |
case-19 | pass→pass | 9,327 | 6,041 | -35% | 1 | 1 | 0% | 1,365 | 1,716 | +26% | 0 | 0 | — |
case-20 | pass→pass | 8,541 | 7,316 | -14% | 1 | 1 | 0% | 1,175 | 1,779 | +51% | 0 | 0 | — |
case-21 | pass→pass | 15,640 | 10,368 | -34% | 1 | 1 | 0% | 2,235 | 2,236 | +0% | 0 | 0 | — |
case-22 | pass→pass | 28,338 | 11,483 | -59% | 1 | 1 | 0% | 3,182 | 2,822 | -11% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +45 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.