Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when deciding what goes in the ACM CoNEXT main body versus the appendix budget versus the artifact — splitting content by decision-criticality so nothing a reviewer needs to accept the paper hides in an appendix or repository, within the acmart limits (long ≤4 appendix pages, short ≤2).
.claude/skills/brycewang-stanford-conext-supplementary/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -15% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -18% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -25% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 70% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 15% | 0% |
Decide where each piece of content lives: the reviewed body, the appendix budget, or the artifact. The governing rule at CoNEXT is decision-criticality — anything a reviewer must read to be convinced the paper should be accepted belongs in the body, not in an appendix a reviewer may skip or an artifact they may not open. CoNEXT's appendix budget is small (long papers ≤4 pages, short papers ≤2), so this is a real allocation problem, not a dumping ground.
| Tier | Belongs in | Test | |---|---|---| | Decision-critical | Main body (≤16 / ≤10 pages) | Would a reviewer's accept/reject flip without it? -> body | | Supporting | Appendix (≤4 / ≤2 pages) | Strengthens or reassures but does not decide -> appendix | | Reproducibility / bulk | Artifact | Needed to reproduce but not to judge -> artifact |
appendix reads as hidden weakness.
scale).
Reviewers are not required to read appendices or run artifacts to reach a decision; if acceptance hinges on it, it is body content.
contribution.
Keep it within the acmart appendix budget and clearly cross-referenced from the body; an appendix the body never points to is wasted.
conext-reproducibility).
judgment.
The artifact supports reproducibility and (if you opted in) badging; it is not a place to smuggle decision-critical evidence past the page limit.
| Misallocation | Why it hurts | Fix | |---|---|---| | Headline result only in the artifact | Reviewers may not open it; reads as hidden | Bring the result into the body | | Central limitation buried in an appendix | Misses the chance to pre-empt the objection | Argue it in the body where the result lives | | Key baseline tuning only in a repo README | Soundness cannot be judged from the paper | Summarize the tuning in the body | | Appendix overflowing the budget | acmart non-compliance / desk risk | Move bulk to the artifact, keep supporting material | | Body padded with material that belongs in the artifact | Wastes pages you need for evidence | Push reproduction bulk to the artifact |
Every tier is a leak surface: an appendix figure showing an internal hostname, or an artifact with commit metadata, breaks anonymity as surely as the body. Run the anonymity sweep (see conext-submission) across the body, appendix, and artifact together.
text[Allocation] body / appendix / artifact for each major piece of content [Decision-critical check] anything acceptance depends on living outside the body? -> move it in [Budget] appendix pages used vs. limit (long ≤4 / short ≤2); acmart compliant? [Cross-refs] appendix + artifact pointed to from the body? yes/no [Anonymity] body + appendix + artifact swept together? yes/no
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 21,949 | 12,264 | -44% | 1 | 1 | 0% | 2,559 | 2,182 | -15% | 0 | 0 | — |
case-02 | fail→fail | 28,408 | 8,540 | -70% | 1 | 1 | 0% | 3,544 | 2,325 | -34% | 0 | 0 | — |
case-03 | fail→pass | 32,904 | 11,309 | -66% | 1 | 1 | 0% | 2,720 | 2,219 | -18% | 0 | 0 | — |
case-04 | pass→pass | 16,078 | 13,150 | -18% | 1 | 1 | 0% | 1,734 | 2,218 | +28% | 0 | 0 | — |
case-05 | pass→pass | 21,130 | 20,707 | -2% | 1 | 1 | 0% | 2,783 | 3,478 | +25% | 0 | 0 | — |
case-06 | pass→pass | 12,632 | 18,688 | +48% | 1 | 1 | 0% | 2,048 | 3,236 | +58% | 0 | 0 | — |
case-07 | fail→pass | 24,691 | 13,673 | -45% | 1 | 1 | 0% | 2,993 | 2,244 | -25% | 0 | 0 | — |
case-08 | fail→pass | 14,709 | 14,267 | -3% | 1 | 1 | 0% | 1,431 | 2,427 | +70% | 0 | 0 | — |
case-09 | fail→pass | 19,524 | 14,531 | -26% | 1 | 1 | 0% | 2,160 | 2,492 | +15% | 0 | 0 | — |
case-10 | fail→pass | 19,416 | 13,093 | -33% | 1 | 1 | 0% | 2,825 | 2,213 | -22% | 0 | 0 | — |
case-11 | fail→pass | 17,386 | 8,445 | -51% | 1 | 1 | 0% | 2,039 | 2,267 | +11% | 0 | 0 | — |
case-12 | pass→pass | 13,828 | 13,861 | +0% | 1 | 1 | 0% | 2,126 | 2,369 | +11% | 0 | 0 | — |
case-13 | fail→fail | 22,916 | 16,152 | -30% | 1 | 1 | 0% | 2,807 | 2,371 | -16% | 0 | 0 | — |
case-14 | fail→fail | 18,529 | 12,503 | -33% | 1 | 1 | 0% | 2,267 | 2,221 | -2% | 0 | 0 | — |
case-15 | pass→pass | 20,103 | 13,839 | -31% | 1 | 1 | 0% | 1,901 | 2,317 | +22% | 0 | 0 | — |
case-16 | pass→pass | 23,721 | 12,848 | -46% | 1 | 1 | 0% | 2,867 | 2,132 | -26% | 0 | 0 | — |
case-17 | fail→pass | 31,412 | 13,966 | -56% | 1 | 1 | 0% | 2,387 | 2,281 | -4% | 0 | 0 | — |
case-18 | fail→pass | 17,552 | 7,745 | -56% | 1 | 1 | 0% | 2,253 | 2,167 | -4% | 0 | 0 | — |
case-19 | pass→pass | 19,556 | 14,420 | -26% | 1 | 1 | 0% | 2,147 | 2,525 | +18% | 0 | 0 | — |
case-20 | fail→pass | 12,587 | 15,167 | +20% | 1 | 1 | 0% | 1,965 | 2,555 | +30% | 0 | 0 | — |
case-21 | fail→pass | 24,571 | 15,178 | -38% | 1 | 1 | 0% | 2,267 | 2,208 | -3% | 0 | 0 | — |
case-22 | fail→pass | 13,486 | 9,138 | -32% | 1 | 1 | 0% | 1,984 | 2,128 | +7% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +55 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.