Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when designing or auditing the empirical work behind a CSCW paper — interview and ethnographic rigor, trace and log analysis, surveys, deployments, and mixed methods — matching each method's own validity standard and the ethics of studying real communities.
.claude/skills/brycewang-stanford-cscw-experiments/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | 40% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 20% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -5% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 47% | 0% |
"Experiments" is the wrong word for most CSCW evidence, and that is the point. The venue is deliberately methods-pluralist: interview studies, ethnography, large-scale trace analysis, surveys, field deployments, controlled experiments, and mixed designs all publish here — each judged by its own tradition's standard of rigor, not by a quantitative default. The commonest reviewing disaster is a paper that borrows a method without its discipline.
| Method | What rigor means here | What reviewers flag | | --- | --- | --- | | Interviews | Purposeful sampling with a rationale; saturation or a defended stopping rule; a described analysis process (coding approach, memoing, disagreement handling) | "We interviewed 12 people and themes emerged" with no analytic trail | | Ethnography / field observation | Duration and depth of engagement; researcher's relationship to the setting; thick description that earns the interpretation | Drive-by observation labeled ethnography | | Trace / log analysis | Construct validity (does the log field measure the practice claimed?); an identification strategy for any causal wording; robustness to platform quirks (bots, deleted content, API sampling) | Correlational results narrated causally; metrics inherited from the platform unexamined | | Surveys | Instrument provenance or validation; sampling frame vs. claimed population; nonresponse handling | Convenience sample generalized to "users" | | Deployments / experiments | Genuine group-level conditions; power analysis where inference is statistical; contamination between conditions addressed | N = groups treated as N = individuals | | Mixed methods | An explicit integration logic — which strand leads, which bounds, where they may disagree | Two mini-studies stapled together, each too thin to stand |
Two pluralism rules cut across all rows:
inter-rater reliability statistics from an interpretivist analysis, or accepting vibes in place of identification from a causal claim, are the same category of error. State which tradition your analysis works in, then meet that tradition fully.
well-chosen moderators can ground a concept; two million log rows cannot rescue a construct that measures the wrong thing.
observe individuals (interviewees, accounts) while claiming about collectives (teams, communities); the analysis section must say how the aggregation is licensed.
how communities were chosen and what variation the set covers — community selection is the qualitative analogue of a sampling frame.
cycles, norm renegotiations). A snapshot design should say what it cannot see.
CSCW evidence usually comes from real communities with stakes in the findings. Reviewers read the ethics description as part of the method:
including whether "public" trace data was treated as fair game and why that is defensible for this community.
communities can be reverse-searched; commit to paraphrase or alteration policies and disclose them.
community can harm it even with every user anonymized (see cscw-artifact-evaluation for release-time handling).
text[Claim] <the group-level finding this study must support> [Tradition] interpretivist / positivist / computational / mixed (lead strand: ___) [Observation] who or what is observed, at what unit and timescale [Aggregation] how individual observations license collective claims [Validity] the ONE threat most likely to sink this design + mitigation [Ethics] consent posture per data source; quote policy; community exposure [Stop rule] what tells you data collection is done
Draft this skeleton before collecting anything; paste the filled version into the methods section as its outline. Under Revise and Resubmit, new data collection is often infeasible — a design that anticipates the obvious objection is the cheapest insurance the journal model offers.
Method norms are stable venue culture; submission-mechanics facts elsewhere in this pack carry the 2026-07-08 access date and should be re-verified independently.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-10 | fail→pass | 15,946 | 15,774 | -1% | 1 | 1 | 0% | 2,550 | 3,572 | +40% | 0 | 0 | — |
case-11 | fail→fail | 14,356 | 14,800 | +3% | 1 | 1 | 0% | 2,485 | 2,705 | +9% | 0 | 0 | — |
case-01 | fail→pass | 24,839 | 20,987 | -16% | 1 | 1 | 0% | 2,797 | 3,354 | +20% | 0 | 0 | — |
case-02 | fail→fail | 32,863 | 25,086 | -24% | 1 | 1 | 0% | 4,178 | 3,845 | -8% | 0 | 0 | — |
case-03 | fail→pass | 44,772 | 38,219 | -15% | 1 | 1 | 0% | 6,512 | 6,209 | -5% | 0 | 0 | — |
case-04 | pass→pass | 23,365 | 22,487 | -4% | 1 | 1 | 0% | 3,244 | 3,867 | +19% | 0 | 0 | — |
case-05 | pass→pass | 20,123 | 21,821 | +8% | 1 | 1 | 0% | 3,157 | 3,911 | +24% | 0 | 0 | — |
case-06 | pass→pass | 47,594 | 17,346 | -64% | 1 | 1 | 0% | 3,007 | 4,130 | +37% | 0 | 0 | — |
case-07 | fail→pass | 25,156 | 20,130 | -20% | 1 | 1 | 0% | 3,374 | 3,274 | -3% | 0 | 0 | — |
case-08 | pass→pass | 18,264 | 17,619 | -4% | 1 | 1 | 0% | 2,221 | 3,457 | +56% | 0 | 0 | — |
case-09 | fail→pass | 21,246 | 12,923 | -39% | 1 | 1 | 0% | 2,163 | 3,170 | +47% | 0 | 0 | — |
case-12 | pass→pass | 21,359 | 19,552 | -8% | 1 | 1 | 0% | 2,724 | 3,350 | +23% | 0 | 0 | — |
case-13 | pass→pass | 22,792 | 19,755 | -13% | 1 | 1 | 0% | 2,948 | 2,980 | +1% | 0 | 0 | — |
case-14 | fail→pass | 20,059 | 22,403 | +12% | 1 | 1 | 0% | 2,122 | 3,210 | +51% | 0 | 0 | — |
case-15 | pass→pass | 14,649 | 19,866 | +36% | 1 | 1 | 0% | 2,231 | 3,040 | +36% | 0 | 0 | — |
case-16 | pass→pass | 15,017 | 18,861 | +26% | 1 | 1 | 0% | 2,221 | 2,969 | +34% | 0 | 0 | — |
case-17 | fail→fail | 16,060 | 19,021 | +18% | 1 | 1 | 0% | 2,414 | 2,969 | +23% | 0 | 0 | — |
case-18 | pass→pass | 29,467 | 13,473 | -54% | 1 | 1 | 0% | 1,921 | 2,411 | +26% | 0 | 0 | — |
case-19 | fail→pass | 18,406 | 14,326 | -22% | 1 | 1 | 0% | 1,981 | 2,392 | +21% | 0 | 0 | — |
case-20 | pass→pass | 20,470 | 7,616 | -63% | 1 | 1 | 0% | 2,512 | 2,045 | -19% | 0 | 0 | — |
case-21 | pass→pass | 22,630 | 21,196 | -6% | 1 | 1 | 0% | 2,790 | 3,502 | +26% | 0 | 0 | — |
case-22 | fail→fail | 19,671 | 17,616 | -10% | 1 | 1 | 0% | 2,144 | 2,794 | +30% | 0 | 0 | — |
case-23 | pass→pass | 26,302 | 22,004 | -16% | 1 | 1 | 0% | 2,842 | 3,386 | +19% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +30 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.