Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when strengthening the transparency of a CSCW paper — auditable qualitative analysis trails, documented trace pipelines, codebooks and instruments, and honest data-availability statements when community and participant data cannot ethically be shared.
.claude/skills/brycewang-stanford-cscw-reproducibility/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -33% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 11% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 19% | 0% |
Reproducibility at CSCW cannot mean "rerun my script, get my table" — most of the venue's evidence is people, and much of it must never leave the research team. The venue's real standard is auditability: a skeptical reader should be able to see how you got from data to claims, and to build on the work, even where they cannot re-execute it. Different strands of a paper owe different transparency debts.
| Evidence strand | Shareable | Auditable instead of shareable | | --- | --- | --- | | Interviews / fieldwork | Interview guide, recruitment text, codebook with definitions and example (paraphrased) excerpts | The analysis trail: coding approach, memo practice, how disagreements were resolved, how themes stabilized | | Trace / log analysis | Pipeline code, query definitions, aggregated datasets, synthetic samples | Exact API/version/date of collection; filtering decisions with counts at each step; bot/deletion handling | | Surveys | Full instrument, scale provenance, analysis scripts | Sampling frame, response/nonresponse accounting | | Deployments | System code or architecture description, condition assignment logic | Site-selection reasoning; what the deployment context makes non-portable | | Statistics anywhere | Analysis scripts keyed to each table/figure | Pre-specification vs. exploration, stated honestly |
You cannot share transcripts; you can share how you thought. The auditable minimum for interpretive work:
inclusion/exclusion notes, and one paraphrased exemplar each. Whether inter-rater statistics belong depends on the tradition; saying which tradition and why is the transparency act.
disconfirming cases forced revisions. Two paragraphs in an appendix outperform a ritual "themes emerged."
participant and context, with the paraphrase/alteration policy stated in the paper.
Platform data rots. Reviewers and future researchers need the ledger even when the data cannot travel:
text[Source] platform, endpoint/API version, collection dates [Scope] query terms / community list / time window, with the WHY [Attrition] rows at each filter step: raw → deduplicated → bot-filtered → analysis set (counts, not adjectives) [Constructs] each analysis variable → the raw field(s) it derives from → the practice it is claimed to measure [Fragility] what breaks if the platform changes (API terms, deletion policy) [Release] what is shared: code / aggregates / synthetic sample / nothing + reason
Write the data statement as a truth-telling exercise, not boilerplate. Three honest shapes:
cannot be redistributed under the platform's terms and our ethics protocol."
are not shareable under the consent participants gave — we chose consent terms that protected candor over shareability, and say so."
pipeline verification."
What never survives review twice (remember the same reviewers return at R&R): "data available upon reasonable request" with no request path, and claims of sharing that the supplement does not actually contain.
For confirmatory quantitative strands, preregistration strengthens the paper — link it anonymized (registries support anonymous view links). Do not force exploratory or interpretive work into a preregistration costume; labeling exploration honestly is the venue's norm.
text[Per strand] shareable artifacts listed and actually present? y/n [Qualitative] codebook + decision log exist? tradition named? y/n [Trace] ledger complete incl. attrition counts? y/n [Statement] availability text matches reality exactly? y/n [Ethics gate] every shared artifact re-checked against consent scope? y/n
Run the gate last and strictly: a transparency package that violates a consent agreement is not a reproducibility win, it is a research-ethics failure that cscw-artifact-evaluation exists to prevent.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 14,932 | 17,181 | +15% | 1 | 1 | 0% | 2,510 | 2,926 | +17% | 0 | 0 | — |
case-02 | fail→pass | 39,077 | 21,368 | -45% | 1 | 1 | 0% | 5,605 | 3,762 | -33% | 0 | 0 | — |
case-03 | pass→pass | 20,070 | 25,071 | +25% | 1 | 1 | 0% | 3,625 | 4,407 | +22% | 0 | 0 | — |
case-04 | pass→pass | 24,861 | 22,687 | -9% | 1 | 1 | 0% | 3,078 | 3,595 | +17% | 0 | 0 | — |
case-05 | pass→pass | 27,303 | 22,893 | -16% | 1 | 1 | 0% | 3,388 | 3,763 | +11% | 0 | 0 | — |
case-06 | pass→pass | 11,449 | 12,119 | +6% | 1 | 1 | 0% | 1,738 | 2,057 | +18% | 0 | 0 | — |
case-07 | fail→pass | 18,641 | 13,180 | -29% | 1 | 1 | 0% | 2,045 | 2,272 | +11% | 0 | 0 | — |
case-08 | fail→pass | 16,581 | 14,056 | -15% | 1 | 1 | 0% | 1,922 | 2,196 | +14% | 0 | 0 | — |
case-09 | fail→pass | 13,959 | 15,933 | +14% | 1 | 1 | 0% | 2,299 | 2,730 | +19% | 0 | 0 | — |
case-10 | fail→pass | 21,952 | 11,161 | -49% | 1 | 1 | 0% | 2,736 | 2,711 | -1% | 0 | 0 | — |
case-11 | fail→pass | 18,288 | 13,401 | -27% | 1 | 1 | 0% | 1,970 | 2,299 | +17% | 0 | 0 | — |
case-12 | pass→pass | 20,494 | 20,750 | +1% | 1 | 1 | 0% | 2,438 | 3,318 | +36% | 0 | 0 | — |
case-13 | pass→pass | 17,731 | 13,635 | -23% | 1 | 1 | 0% | 1,924 | 2,308 | +20% | 0 | 0 | — |
case-14 | fail→pass | 21,598 | 16,101 | -25% | 1 | 1 | 0% | 2,858 | 2,566 | -10% | 0 | 0 | — |
case-15 | fail→pass | 20,450 | 16,372 | -20% | 1 | 1 | 0% | 2,531 | 2,760 | +9% | 0 | 0 | — |
case-16 | fail→pass | 20,605 | 17,647 | -14% | 1 | 1 | 0% | 2,082 | 2,762 | +33% | 0 | 0 | — |
case-17 | pass→pass | 15,061 | 13,666 | -9% | 1 | 1 | 0% | 1,696 | 2,171 | +28% | 0 | 0 | — |
case-18 | pass→pass | 18,798 | 10,943 | -42% | 1 | 1 | 0% | 2,088 | 1,770 | -15% | 0 | 0 | — |
case-19 | pass→pass | 15,544 | 12,410 | -20% | 1 | 1 | 0% | 2,593 | 3,033 | +17% | 0 | 0 | — |
case-20 | fail→pass | 12,117 | 8,753 | -28% | 1 | 1 | 0% | 1,940 | 1,665 | -14% | 0 | 0 | — |
case-21 | fail→pass | 16,350 | 15,628 | -4% | 1 | 1 | 0% | 2,320 | 2,597 | +12% | 0 | 0 | — |
case-22 | pass→pass | 11,350 | 16,927 | +49% | 1 | 1 | 0% | 1,844 | 2,557 | +39% | 0 | 0 | — |
case-23 | fail→fail | 16,189 | 14,415 | -11% | 1 | 1 | 0% | 2,829 | 2,783 | -2% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +52 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.