Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when packaging AAAI code, data, multimedia appendices, technical appendices, reproducibility evidence, and post-acceptance artifact releases without violating double-blind or immutable-supplement rules.
.claude/skills/brycewang-stanford-aaai-artifact-evaluation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -30% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -36% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 9% | 0% |
Use this to prepare artifacts that reviewers can use to assess reproducibility. AAAI supplementary material is part of the submission record; after review starts, do not assume it can be updated.
experiments.
commands, seeds, expected outputs, and runtime.
rebuttal.
AAAI does not run a separate badged artifact-evaluation committee the way some systems venues do; the same broad-AI reviewer who scores the paper also inspects whatever supplement you attach. That reviewer may be a planning, knowledge-representation, or constraint-satisfaction specialist rather than a deep-learning engineer, so the artifact has to be legible without insider tooling. Optimize for a reviewer who skims, not one who will spend an afternoon configuring a cluster.
| Reviewer action | Passes | Fails | | --- | --- | --- | | Opens the ZIP | sane tree, top README | nested archives, 0-byte files | | Reads appendix | maps to numbered claims | contradicts the paper | | Tries one command | reproduces one headline number | needs private data or credentials | | Scans for identity | nothing reveals authors | Git logs or home paths leak |
Because clearly-below-bar papers can be cut before author feedback, a supplement that looks thin or unrunnable is a cheap reason to summary-reject. Avoid these:
A constraint-solving paper claims a 30% node-expansion reduction. The team ships a large ZIP of raw solver logs but no driver script. The reproduction path is empty, so artifact status is "risky"; the fix is a small run_main.py that regenerates Table 2 from seeds, a trimmed log sample, and a license for the benchmark instances. The raw dump moves to the post-acceptance release.
text[Artifact status] complete / partial / risky / unavailable [Submitted files] technical appendix / multimedia appendix / code-data ZIP [Reviewer reproduction path] <commands and expected output> [Anonymity risks] <metadata, links, paths, logs> [Missing items] <data, code, seeds, licenses, hardware>
Other measured skills in the registry, with their headline benchmark lift.