Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Gather evidence for causal claims
.claude/skills/yogsoth-ai-evidence-collection/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 52% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 299% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 64% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -25% | 0% |
Systematically gather and attach evidence pages to the causal claims in the model, creating supported_by edges for confirming evidence and contradicts edges for disconfirming evidence. A causal model without evidence provenance is speculation; this strategy converts it into a structured, auditable knowledge artifact.
CC should treat evidence collection as adversarial by default: for every supported_by edge added, actively search for contradicting evidence before moving on. The contradicts edges are as important as the supported_by edges — a model that only records confirming evidence is biased. Evidence pages must record the source, study design (where applicable), and a brief assessment of quality. Quantity matters less than coverage across independent sources.
| Metric | S | M | L | |--------|---|---|---| | Evidence pages created | 8 | 20 | 45 | | supported_by edges | 10 | 25 | 50 | | contradicts edges flagged | 2 | 5 | 10 |
| Metric | Target | Current | Status |
|--------------------------|--------|---------|--------|
| Evidence pages created | S:8 / M:20 / L:45 | 0 | ⬜ |
| supported_by edges | S:10 / M:25 / L:50 | 0 | ⬜ |
| contradicts edges flagged | S:2 / M:5 / L:10 | 0 | ⬜ |<HARD-GATE> Cannot exit until 80% of budget met. Print state ledger before each iteration decision. </HARD-GATE>
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | counterfactual-reasoning | Tactic for reasoning about what would happen if variables were different — supports causal identification and intervention analysis. | | evidence-weighing | Tactic for assessing the strength and relevance of evidence for causal claims — distinguishes correlation from causation. | | feedback-loop-detection | Tactic for identifying circular causation — detect feedback loops, classify as reinforcing or balancing, document loop structure. |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | evidence-linking | SOP for linking evidence pages to causal claims — creates supported_by or contradicts edges. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-21 | pass→pass | 13,354 | 30,469 | +128% | 1 | 1 | 0% | 1,972 | 5,782 | +193% | 0 | 0 | — |
case-01 | fail→pass | 27,341 | 32,369 | +18% | 1 | 1 | 0% | 4,500 | 6,855 | +52% | 0 | 0 | — |
case-02 | fail→pass | 36,753 | 34,288 | -7% | 1 | 1 | 0% | 6,212 | 6,851 | +10% | 0 | 0 | — |
case-03 | fail→fail | 31,527 | 37,018 | +17% | 1 | 1 | 0% | 5,347 | 6,852 | +28% | 0 | 0 | — |
case-04 | pass→fail | 19,668 | 24,006 | +22% | 1 | 1 | 0% | 3,677 | 4,672 | +27% | 0 | 0 | — |
case-05 | fail→fail | 3,305 | 9,787 | +196% | 1 | 1 | 0% | 499 | 2,197 | +340% | 0 | 0 | — |
case-06 | fail→pass | 9,600 | 27,877 | +190% | 1 | 1 | 0% | 1,706 | 6,810 | +299% | 0 | 0 | — |
case-07 | fail→pass | 14,207 | 15,738 | +11% | 1 | 1 | 0% | 1,971 | 3,228 | +64% | 0 | 0 | — |
case-08 | fail→pass | 14,385 | 6,811 | -53% | 1 | 1 | 0% | 2,308 | 1,739 | -25% | 0 | 0 | — |
case-09 | fail→pass | 14,848 | 5,228 | -65% | 1 | 1 | 0% | 2,172 | 1,466 | -33% | 0 | 0 | — |
case-10 | pass→pass | 15,041 | 27,720 | +84% | 1 | 1 | 0% | 2,567 | 5,693 | +122% | 0 | 0 | — |
case-11 | fail→pass | 14,615 | 10,870 | -26% | 1 | 1 | 0% | 2,172 | 2,337 | +8% | 0 | 0 | — |
case-12 | pass→pass | 12,191 | 14,885 | +22% | 1 | 1 | 0% | 2,013 | 2,999 | +49% | 0 | 0 | — |
case-13 | pass→pass | 12,476 | 10,006 | -20% | 1 | 1 | 0% | 2,138 | 2,327 | +9% | 0 | 0 | — |
case-14 | pass→pass | 10,569 | 12,917 | +22% | 1 | 1 | 0% | 1,625 | 2,652 | +63% | 0 | 0 | — |
case-15 | fail→pass | 9,528 | 2,173 | -77% | 1 | 1 | 0% | 1,365 | 1,005 | -26% | 0 | 0 | — |
case-16 | fail→pass | 9,605 | 2,174 | -77% | 1 | 1 | 0% | 1,387 | 948 | -32% | 0 | 0 | — |
case-17 | fail→pass | 9,032 | 2,319 | -74% | 1 | 1 | 0% | 1,417 | 982 | -31% | 0 | 0 | — |
case-18 | pass→pass | 11,712 | 1,923 | -84% | 1 | 1 | 0% | 1,647 | 866 | -47% | 0 | 0 | — |
case-19 | fail→pass | 5,568 | 1,718 | -69% | 1 | 1 | 0% | 790 | 857 | +8% | 0 | 0 | — |
case-20 | fail→pass | 8,788 | 6,976 | -21% | 1 | 1 | 0% | 1,385 | 1,760 | +27% | 0 | 0 | — |
case-22 | fail→pass | 13,248 | 6,906 | -48% | 1 | 1 | 0% | 2,218 | 1,801 | -19% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +55 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.