Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when designing or auditing AISTATS experiments, simulations, baselines, statistical tests, uncertainty estimates, ablations, random seeds, hyperparameters, compute, dataset handling, and claim-to-evidence fit, with emphasis on experiments that validate theorems rather than chase leaderboards.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -11% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 26% | 0% |
Use this before submission when the empirical or simulation story is not yet locked.
show practical relevance.
intervals, paired tests, or bootstrap intervals when appropriate.
settings, selection criteria, random seeds, hardware, software versions, and runtime.
theoretical assumptions and empirical setup.
simulation confirming a predicted rate outweighs five extra benchmark datasets.
they are deliberately violated, and a real-data study showing practical behavior.
dimension, noise level — matches the asymptotic regime of the theorems. A bound proven as n grows but tested only at n = 500 invites the question of relevance.
| Theoretical claim | Matching experiment | Reject pattern avoided | |---|---|---| | Convergence rate in n | Log-log error versus n with fitted slope | "Rates asserted but never plotted" | | Confidence-interval coverage | Empirical coverage across many replications | "Nominal 95 percent never verified" | | Regret bound | Cumulative regret versus horizon, with the bound curve overlaid | "Bound and trajectory never compared" | | Robustness to misspecification | Violation-severity sweep | "Guarantees hold under assumptions the experiments quietly break" |
Suppose the paper proves finite-sample type-I error control under a boundedness assumption. The matching plan: simulate under the null at several sample sizes to verify size, sweep dependence strength for power curves, then inject heavy-tailed noise that breaks boundedness to map degradation — every panel tied to a numbered theorem or remark.
are standard errors, confidence intervals, or quantiles.
text[Experiment readiness] strong / adequate / weak [Claim -> evidence map] <claim: table/figure/simulation> [Missing statistical evidence] <uncertainty/test/seed/baseline> [Reproducibility gaps] <hyperparameters/compute/data/code> [Decision-critical next run] <one experiment or simulation>
Other measured skills in the registry, with their headline benchmark lift.