Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use in Stage 08 (Dissemination) when assembling the release or submission bundle — auditing whether the run's code, data, results and figures are actually reproducible by someone else, writing the readiness checklist and threats-to-validity notes, or deciding what has to be disclosed as not verified.
.claude/skills/tangxiangru-reproducibility-check/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -23% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 75% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -5% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -44% | 0% |
Stage 08 must write review and readiness artifacts under workspace/reviews/. The purpose is to state honestly what a third party could reproduce from this run's directory alone — not to assert that everything is fine.
The output of this check is a record of what was verified, what was not verified, and what is known not to reproduce. All three categories are legitimate. Only a fourth is not: recording an unchecked item as verified.
Work from runs/<run_id>/ and assume no access to the conversation, the operator's memory, or the environment the run happened in.
| Question | Where the answer must live | | --- | --- | | What was the research goal? | user_input.txt | | What was actually run? | workspace/code/ — scripts, not descriptions | | What data went in? | workspace/data/, with provenance for anything downloaded | | What came out? | workspace/results/, indexed by experiment_manifest.json | | How do the figures relate to the results? | a script under workspace/code/ that regenerates them | | What decisions were made and why? | the decision ledger in the stage summaries | | Which stages actually completed? | run_manifest.json — see below |
run_manifest.json distinguishes stages that were approved from stages that were skipped. A skipped stage has skipped: true and a skip_kind of human or auto; an auto skip means the stage exhausted its retry budget in an unattended run with nobody in the loop, and its work was never done.
A readiness review that reports a run as complete when a stage was auto-skipped is wrong in the most damaging direction. List every skipped stage in the readiness artifact, with its skip_reason, and say what downstream claim is weakened by the gap.
For each item, record verified, not_verified, or fails, with a one-line reason. Never leave an item without evidence for its status.
dependencies? A script that imports a package the run never declares is not reproducible.
say by how much.
workspace/data/, where did it comefrom, and can someone else obtain it? "Generated synthetically" is a complete answer only if the generator is in workspace/code/.
there a gap where a manual step happened?
in the run? citation_verification.json covers external claims; this covers the run's own.
venue-checklist skill.
Write these in the paper's own voice, not as a disclaimer appendix nobody reads. The ones that matter most in an AutoR run:
split.
manifest against what Stage 06 actually measured, and flag anything the manuscript asserts more strongly than the analysis supports.
workspace/artifacts/ contains what the checklist says itcontains.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 5,395 | 14,129 | +162% | 1 | 1 | 0% | 283 | 1,178 | +316% | 0 | 0 | — |
case-02 | fail→fail | 14,061 | 15,205 | +8% | 1 | 1 | 0% | 228 | 1,245 | +446% | 0 | 0 | — |
case-03 | fail→fail | 14,609 | 11,432 | -22% | 1 | 1 | 0% | 284 | 1,241 | +337% | 0 | 0 | — |
case-04 | fail→pass | 11,443 | 3,630 | -68% | 1 | 1 | 0% | 2,044 | 1,579 | -23% | 0 | 0 | — |
case-05 | fail→pass | 12,011 | 2,883 | -76% | 1 | 1 | 0% | 1,822 | 1,514 | -17% | 0 | 0 | — |
case-06 | fail→pass | 12,146 | 6,344 | -48% | 1 | 1 | 0% | 1,261 | 2,213 | +75% | 0 | 0 | — |
case-07 | pass→pass | 16,312 | 10,975 | -33% | 1 | 1 | 0% | 1,783 | 2,113 | +19% | 0 | 0 | — |
case-08 | fail→pass | 12,372 | 6,863 | -45% | 1 | 1 | 0% | 1,295 | 1,224 | -5% | 0 | 0 | — |
case-09 | fail→pass | 13,529 | 1,855 | -86% | 1 | 1 | 0% | 2,219 | 1,249 | -44% | 0 | 0 | — |
case-10 | fail→pass | 13,651 | 5,273 | -61% | 1 | 1 | 0% | 2,188 | 1,810 | -17% | 0 | 0 | — |
case-11 | pass→pass | 10,451 | 17,294 | +65% | 1 | 1 | 0% | 1,764 | 2,302 | +30% | 0 | 0 | — |
case-12 | fail→pass | 9,528 | 8,800 | -8% | 1 | 1 | 0% | 1,463 | 1,531 | +5% | 0 | 0 | — |
case-13 | pass→pass | 14,229 | 11,312 | -21% | 1 | 1 | 0% | 1,427 | 2,006 | +41% | 0 | 0 | — |
case-14 | fail→pass | 14,462 | 9,609 | -34% | 1 | 1 | 0% | 1,360 | 1,662 | +22% | 0 | 0 | — |
case-15 | fail→pass | 27,821 | 7,838 | -72% | 1 | 1 | 0% | 1,589 | 1,398 | -12% | 0 | 0 | — |
case-16 | fail→pass | 17,439 | 3,858 | -78% | 1 | 1 | 0% | 1,965 | 1,606 | -18% | 0 | 0 | — |
case-17 | pass→pass | 12,404 | 6,447 | -48% | 1 | 1 | 0% | 2,115 | 2,091 | -1% | 0 | 0 | — |
case-18 | pass→pass | 10,981 | 9,762 | -11% | 1 | 1 | 0% | 1,847 | 2,497 | +35% | 0 | 0 | — |
case-19 | pass→pass | 7,351 | 8,497 | +16% | 1 | 1 | 0% | 1,023 | 1,572 | +54% | 0 | 0 | — |
case-20 | pass→pass | 7,980 | 8,883 | +11% | 1 | 1 | 0% | 1,398 | 1,606 | +15% | 0 | 0 | — |
case-21 | pass→pass | 15,672 | 37,335 | +138% | 1 | 1 | 0% | 2,276 | 4,392 | +93% | 0 | 0 | — |
case-22 | pass→fail | 17,495 | 10,433 | -40% | 1 | 1 | 0% | 2,815 | 1,205 | -57% | 0 | 0 | — |
case-23 | fail→fail | 7,461 | 5,454 | -27% | 1 | 1 | 0% | 1,189 | 1,198 | +1% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 17 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +39 percentage points is the difference between those two pass rates over the 17 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.