Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when designing or auditing AAAI experiments for the broad-AI program committee, including baselines, ablations, statistical significance, robustness, human evaluation, AI-for-Social-Impact and alignment/safety evidence, compute and cost reporting, and reproducibility-checklist alignment for Phase-1 survival.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -6% | 0% |
Use this before submission to ensure empirical evidence supports the AI contribution. AAAI reviewers may come from adjacent AI subfields, so experiments must be interpretable beyond one benchmark community.
when relevant.
and ethics/IRB status.
Build this table before adding new experiments. It keeps the AAAI evidence package aligned with the main text and with the reproducibility checklist.
| Manuscript claim | Required evidence | Phase-1 risk if missing | Checklist hook | | --- | --- | --- | --- | | New AI capability | benchmark + qualitative failure cases | broad reviewer sees only engineering | datasets, metrics, baselines | | Better mechanism | single-factor ablations | gain looks like tuning luck | ablation and hyperparameter answers | | Robust deployment | shift / seed / subgroup stress test | result seems brittle | variance, compute, environment | | Social-impact or safety claim | stakeholder, harm, and misuse analysis | ethical claim looks asserted | ethics, limitations, data access |
For each row, mark ready / weak / missing and name the fastest fix that can be run before the supplementary-material deadline. Do not leave a claim in the abstract if its evidence row is weak.
risk mitigation, and scope.
result provenance, seeds, data splits, and limits machine-readable enough that a human SPC/AC can quickly audit them.
Before submission, decide which experiments would be impossible to add later under AAAI's rebuttal constraints: missing baselines, missing seeds, missing supplement files, or missing reproducibility checklist answers. Treat those as pre-submission blockers, not rebuttal TODOs. The author response can explain and clarify submitted evidence; it should not depend on new results, URLs, or repaired supplementary files.
Because an AAAI reviewer from an adjacent subfield must trust your numbers quickly, classify each experimental block by how much weight it can bear and what would strengthen it.
| Block | Carries the claim when | Reviewer doubt | Cheap reinforcement | | --- | --- | --- | --- | | Headline benchmark | beats tuned recent baselines | "lucky seed" | seeds, variance bars | | Ablation | isolates one mechanism | "joint removal" | single-factor toggles | | Robustness | holds across split/shift | "one setting" | extra split or perturbation | | Human eval | protocol is documented | "rater bias" | IRB note, inter-rater agreement |
insight.
A planning paper reports a single-seed win on one domain. Audit: the headline block "needs robustness" and "needs variance", so the fix before the deadline is five seeds with confidence intervals plus one extra IPC-style domain. Because new results cannot rescue this in rebuttal, the team runs both before submission and aligns the checklist's seed answer to the supplement.
text[Claim] <paper claim> [Evidence status] sufficient / needs baseline / needs ablation / needs robustness / unclear [Fairness issue] <compute, tuning, data, prompt, metric, human eval> [Checklist dependency] <what checklist answer this supports> [Pre-rebuttal blockers] <missing evidence that must be run before submission> [Fast fix] <experiment or analysis feasible before deadline>
Other measured skills in the registry, with their headline benchmark lift.