Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when building the reproducibility story for an ACM CoNEXT paper — pinned traces and configs, a runnable artifact, an honest data-availability posture, and the one-page artifact description the CoNEXT reproducibility committee needs — remembering that the ACM badge opt-in is due before the submission deadline.
.claude/skills/brycewang-stanford-conext-reproducibility/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -26% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 11% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 23% | 0% |
Build the reproducible-networking story alongside the experiments, not after acceptance. CoNEXT runs a dedicated reproducibility committee that awards optional ACM badges, and the single most missed rule is that badge eligibility requires opting in before the paper submission deadline — you cannot bolt it on later. Even for authors who skip badging, a reproducible artifact strengthens a double-anonymous, one-shot-revision review.
text[Before the submission deadline] OPT IN for ACM badging (required for eligibility) [At submission] anonymized, runnable artifact referenced from the paper [Within ~1 week of acceptance] send a ONE-PAGE artifact description to the reproducibility committee with pointers to code and other artifacts [Camera-ready / post-accept] committee evaluates for Available / Functional / Reusable / Reproduced badges (see conext-artifact-evaluation)
Miss the opt-in and the strongest artifact in the world cannot earn a badge this cycle.
Networking reproducibility is harder than "here is the code," because the result depends on an environment:
extraction scripts that turn raw captures into the paper's inputs.
representative example.
a testbed; container/VM images where feasible.
regenerate rather than being hand-copied.
You cannot reconstruct provenance after the campaign ends:
reproducible.
that needs live API calls re-samples rather than reproduces.
Heritage) and an open license.
can (aggregate data, synthetic traces, the analysis pipeline). "Available upon request" reads as a scored weakness, not a neutral placeholder.
reviewer or the committee will check.
internal hostnames, and owner-identifying paths.
and IP ranges you own.
After acceptance, the committee wants a one-page description that lets an evaluator start quickly:
networking artifacts need specific switches/NICs — flag this early).
text[Opt-in] badge opt-in done BEFORE the submission deadline? (if badging) yes/no [Traces] captures + vantage points + dates + extraction scripts present? yes/no [Configs] exact per-figure configs and topology recorded? yes/no [Environment] hardware/firmware/OS versions pinned; image where feasible? yes/no [Pipeline] raw -> figure regenerates by script? yes/no [Availability] honest statement; DOI archive or a documented reason not to release? yes/no [Anonymity] artifact re-hosted anonymously; metadata scrubbed? yes/no
text[Reproducibility status] ready / gaps [Badge intent] opt-in before submission? yes/no/n-a [Provenance] traces/configs/environment pinned at collection time [Availability] what is released, where (DOI), and any documented restriction [Anonymity] artifact anonymized for double-anonymous review [One-pager] artifact description drafted for the committee (post-accept)
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 80,683 | 21,903 | -73% | 1 | 1 | 0% | 5,155 | 3,819 | -26% | 0 | 0 | — |
case-02 | fail→fail | 42,175 | 19,706 | -53% | 1 | 1 | 0% | 5,842 | 3,397 | -42% | 0 | 0 | — |
case-03 | fail→pass | 35,198 | 22,824 | -35% | 1 | 1 | 0% | 4,431 | 3,714 | -16% | 0 | 0 | — |
case-04 | fail→pass | 17,583 | 11,937 | -32% | 1 | 1 | 0% | 1,935 | 2,141 | +11% | 0 | 0 | — |
case-05 | pass→pass | 17,873 | 9,276 | -48% | 1 | 1 | 0% | 1,669 | 2,545 | +52% | 0 | 0 | — |
case-06 | fail→pass | 20,647 | 16,351 | -21% | 1 | 1 | 0% | 2,267 | 2,774 | +22% | 0 | 0 | — |
case-07 | pass→pass | 15,312 | 11,241 | -27% | 1 | 1 | 0% | 2,287 | 2,730 | +19% | 0 | 0 | — |
case-08 | fail→pass | 20,065 | 16,505 | -18% | 1 | 1 | 0% | 2,295 | 2,813 | +23% | 0 | 0 | — |
case-09 | fail→fail | 19,518 | 18,038 | -8% | 1 | 1 | 0% | 2,126 | 2,928 | +38% | 0 | 0 | — |
case-10 | pass→pass | 13,946 | 15,281 | +10% | 1 | 1 | 0% | 2,136 | 2,800 | +31% | 0 | 0 | — |
case-11 | fail→pass | 20,900 | 14,296 | -32% | 1 | 1 | 0% | 1,759 | 2,533 | +44% | 0 | 0 | — |
case-12 | pass→pass | 20,205 | 18,370 | -9% | 1 | 1 | 0% | 2,264 | 2,692 | +19% | 0 | 0 | — |
case-13 | fail→fail | 21,859 | 19,334 | -12% | 1 | 1 | 0% | 3,580 | 3,179 | -11% | 0 | 0 | — |
case-14 | fail→pass | 17,609 | 15,726 | -11% | 1 | 1 | 0% | 1,860 | 2,632 | +42% | 0 | 0 | — |
case-15 | fail→pass | 13,980 | 9,613 | -31% | 1 | 1 | 0% | 1,353 | 1,830 | +35% | 0 | 0 | — |
case-16 | pass→pass | 23,249 | 12,281 | -47% | 1 | 1 | 0% | 2,882 | 3,014 | +5% | 0 | 0 | — |
case-17 | pass→pass | 12,273 | 7,739 | -37% | 1 | 1 | 0% | 1,107 | 1,529 | +38% | 0 | 0 | — |
case-18 | pass→pass | 11,650 | 10,071 | -14% | 1 | 1 | 0% | 1,632 | 2,338 | +43% | 0 | 0 | — |
case-19 | fail→pass | 23,452 | 15,648 | -33% | 1 | 1 | 0% | 2,363 | 2,417 | +2% | 0 | 0 | — |
case-20 | pass→pass | 18,990 | 16,517 | -13% | 1 | 1 | 0% | 1,591 | 2,765 | +74% | 0 | 0 | — |
case-21 | pass→pass | 26,609 | 16,217 | -39% | 1 | 1 | 0% | 2,672 | 3,522 | +32% | 0 | 0 | — |
case-22 | fail→pass | 18,157 | 25,758 | +42% | 1 | 1 | 0% | 2,125 | 3,763 | +77% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +45 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.