Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when packaging what stands behind a CSCW paper — systems, analysis pipelines, codebooks, instruments, datasets from real communities — for review-time scrutiny and post-acceptance release, where community-data ethics constrain release more than any badge checklist.
.claude/skills/brycewang-stanford-cscw-artifact-evaluation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 46% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 43% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 26% | 0% |
CSCW has no artifact-evaluation committee stamping badges; the "evaluators" of your artifacts are the reviewers deciding whether to trust your methods, and later the researchers and the studied communities themselves who encounter the released materials. That second audience is the venue's distinctive constraint: a CSCW artifact release is an act with consequences for real people who never signed your consent form.
| Artifact | Review-time form | Post-acceptance form | Governing constraint | | --- | --- | --- | --- | | Study instruments (guides, surveys) | Anonymized in supplement | Public archive | Institution redaction only | | Codebooks | Anonymized in supplement | Public archive | Paraphrased exemplars only | | Analysis code | Anonymized archive | Public repository + DOI archive | Strip identity from history | | Built systems / prototypes | Screenshots/video + code where feasible | Open source if maintainable | Third-party assets, API keys | | Trace datasets | Aggregates or synthetic sample | Aggregates; raw only if terms + ethics allow | Platform ToS + re-identification risk | | Interview/field data | Never | Almost never | Consent scope is absolute |
Before releasing any dataset or exemplar text derived from an online community, run the adversary exercise — assume a motivated actor (a journalist, a harasser, a platform admin, a community member with a grudge):
and land on the original post? If yes, paraphrase or drop it.
be joined against public APIs or archives to recover usernames? Coarsen until the join fails.
member) expose it to raids, deplatforming, or press attention it has not chosen? If the community is small or marginalized, community-level anonymity is part of the ethical contract — name it only with its agreement.
would this release be safe under a hostile future owner of the platform's data?
Document the outcome in a short release memo; it becomes the honest core of the paper's data-availability statement (cscw-reproducibility).
textrelease/ ├── README.md # what claims each artifact supports; setup in ≤ 5 steps ├── LICENSE # code license + data terms, which may differ ├── instruments/ # guides, surveys, recruitment text (redacted) ├── codebook/ # definitions + paraphrased exemplars ├── pipeline/ # collection + analysis code, config, environment spec ├── data/ │ ├── aggregates/ # tables behind each figure │ └── synthetic/ # optional: structure-preserving fake sample └── ETHICS.md # consent scope, what is withheld and WHY, contact path
ETHICS.md is the CSCW-specific file: it tells reusers what they may not do (re-identify, recontact, join against other datasets) and why the withheld parts are withheld. A release that explains its own limits earns more trust than a maximal dump.
archive paths, and hosting URLs (use anonymized repository services, not a lab server whose domain resolves the institution).
stand if a reviewer opens nothing.
one source tree; the deanonymization diff at acceptance should be mechanical (cscw-camera-ready).
text[Claims map] each artifact → the claim it supports (no orphan uploads) [Adversary test] quote-search / join / community-exposure / future-context: pass? [Consent gate] every released item inside consent + ToS scope? y/n [Two builds] review build anonymous; release build attributed; same tree? y/n [ETHICS.md] withholdings explained, reuse limits stated? y/n
No CSCW badge program existed at the 2026-07-08 check; if one appears in a future cycle, its checklist supplements — never replaces — the community-protection tests above.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 30,901 | 21,338 | -31% | 1 | 1 | 0% | 4,063 | 3,693 | -9% | 0 | 0 | — |
case-02 | fail→pass | 24,707 | 30,119 | +22% | 1 | 1 | 0% | 3,482 | 5,082 | +46% | 0 | 0 | — |
case-03 | fail→pass | 21,790 | 18,565 | -15% | 1 | 1 | 0% | 3,052 | 4,363 | +43% | 0 | 0 | — |
case-04 | fail→pass | 21,329 | 17,257 | -19% | 1 | 1 | 0% | 2,582 | 2,944 | +14% | 0 | 0 | — |
case-05 | fail→pass | 12,291 | 10,625 | -14% | 1 | 1 | 0% | 2,042 | 2,577 | +26% | 0 | 0 | — |
case-06 | pass→pass | 16,509 | 17,605 | +7% | 1 | 1 | 0% | 2,319 | 2,896 | +25% | 0 | 0 | — |
case-07 | pass→pass | 19,300 | 16,644 | -14% | 1 | 1 | 0% | 2,140 | 2,874 | +34% | 0 | 0 | — |
case-08 | pass→pass | 19,975 | 20,695 | +4% | 1 | 1 | 0% | 2,538 | 3,107 | +22% | 0 | 0 | — |
case-09 | fail→pass | 14,278 | 17,279 | +21% | 1 | 1 | 0% | 2,094 | 2,769 | +32% | 0 | 0 | — |
case-10 | pass→pass | 19,580 | 11,093 | -43% | 1 | 1 | 0% | 2,198 | 2,379 | +8% | 0 | 0 | — |
case-11 | pass→pass | 21,006 | 16,464 | -22% | 1 | 1 | 0% | 2,284 | 2,612 | +14% | 0 | 0 | — |
case-12 | fail→pass | 16,587 | 11,252 | -32% | 1 | 1 | 0% | 1,842 | 2,068 | +12% | 0 | 0 | — |
case-13 | fail→pass | 22,688 | 14,989 | -34% | 1 | 1 | 0% | 2,263 | 2,466 | +9% | 0 | 0 | — |
case-14 | pass→pass | 17,321 | 15,267 | -12% | 1 | 1 | 0% | 1,760 | 2,648 | +50% | 0 | 0 | — |
case-15 | pass→pass | 14,451 | 12,726 | -12% | 1 | 1 | 0% | 1,851 | 2,109 | +14% | 0 | 0 | — |
case-16 | pass→pass | 21,008 | 18,754 | -11% | 1 | 1 | 0% | 2,478 | 3,100 | +25% | 0 | 0 | — |
case-17 | fail→pass | 17,200 | 17,523 | +2% | 1 | 1 | 0% | 1,915 | 2,825 | +48% | 0 | 0 | — |
case-18 | pass→pass | 16,799 | 11,960 | -29% | 1 | 1 | 0% | 1,809 | 2,100 | +16% | 0 | 0 | — |
case-19 | pass→pass | 28,660 | 24,132 | -16% | 1 | 1 | 0% | 3,269 | 4,257 | +30% | 0 | 0 | — |
case-20 | pass→pass | 20,812 | 13,389 | -36% | 1 | 1 | 0% | 2,584 | 3,108 | +20% | 0 | 0 | — |
case-21 | pass→pass | 12,036 | 13,106 | +9% | 1 | 1 | 0% | 1,973 | 3,240 | +64% | 0 | 0 | — |
case-22 | pass→pass | 12,757 | 14,796 | +16% | 1 | 1 | 0% | 1,997 | 2,513 | +26% | 0 | 0 | — |
case-23 | pass→pass | 24,125 | 18,693 | -23% | 1 | 1 | 0% | 2,866 | 2,971 | +4% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +39 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.