Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when strengthening AISTATS reproducibility evidence, including the official reproducibility checklist, statistical assumptions, proofs, datasets, hyperparameters, random seeds, compute, uncertainty estimates, baselines, code/data release statements, and checklist-to-claim consistency audits.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 46% | 0% |
| case-17 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 3% | 0% |
| case-13 | ✓→✗ | ▼ Worse | 3% | 0% |
| case-09 | ✓→✓ | = Same ✓ | 12% | 0% |
Use this before submission and again before camera-ready. Reopen the current CFP and OpenReview forms to confirm whether a reproducibility checklist is required.
location in the paper, appendix, supplement, or artifact package.
failure modes clearly enough for statistical readers.
hyperparameter ranges, final selected settings, seeds, repeated runs, compute, and runtime.
intervals, paired tests, bootstrap intervals, or repeated trials as appropriate.
in principle.
paper are review-risk multipliers.
| Checklist item | Pure-theory answer | Theory-plus-experiments answer | |---|---|---| | Code availability | NA only if there is literally no computation | Anonymous archive, or an honest stated reason | | Assumptions stated | Every theorem lists its conditions inline | Plus a note on which experiments satisfy them | | Error bars | NA for deterministic results | Required for every stochastic figure and table | | Compute resources | NA | Hardware, runtime, and total number of runs |
Marking NA on an item the paper actually triggers is a recognizable AISTATS red flag, because reviewers cross-check checklist answers against the PDF and read contradictions as carelessness about the rest of the paper.
Consider a submission proving posterior contraction rates for a Bayesian nonparametric model, validated by MCMC simulation. Its reproducibility spine: prior hyperparameters and their selection rule, chain length, burn-in, convergence diagnostics, replication seeds, and a statement of which contraction-theorem conditions the simulated model satisfies — plus one honest sentence about the condition it does not.
For AISTATS, simulations should be turnkey because statistician reviewers actually rerun them; large real-data pipelines may stay scripted with deviations documented. Stating the achieved level honestly beats overpromising turnkey behavior that fails on a clean machine.
text[Claim inventory] <claim -> evidence location> [Checklist status] complete / inconsistent / missing [Statistical reproducibility gaps] <assumptions/seeds/uncertainty/hyperparameters/compute> [Paper fixes] <must appear in main PDF> [Supplement fixes] <appendix or artifact additions>
Other measured skills in the registry, with their headline benchmark lift.