Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Validate causal model consistency
.claude/skills/yogsoth-ai-model-validation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-19 | ✗→✓ | ▲ Improved | 84% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 573% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 26% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 33% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 48% | 0% |
Audit the completed causal model for internal consistency: detect and resolve cycles that should not exist, surface contradictions between evidence and claimed edges, and recalibrate confidence scores where the evidence base has shifted. This strategy is the final quality gate before the model is used for inference or reporting.
CC must approach validation as a skeptic, not a defender. The goal is to find problems, not confirm that the model is fine. Cycles in a directed acyclic graph (DAG) are structural errors unless explicitly modeled as feedback loops with documentation. Contradictions between evidence pages and mechanism edges indicate that either the edge or the evidence assessment is wrong — both must be re-examined. Confidence recalibrations should propagate: if a key mechanism edge loses confidence, all downstream claims that depend on it should be flagged for review as well.
| Metric | S | M | L | |--------|---|---|---| | Cycles checked | 3 | 8 | 15 | | Contradictions resolved | 1 | 3 | 6 | | Confidence recalibrations | 3 | 8 | 15 |
| Metric | Target | Current | Status |
|---------------------------|--------|---------|--------|
| Cycles checked | S:3 / M:8 / L:15 | 0 | ⬜ |
| Contradictions resolved | S:1 / M:3 / L:6 | 0 | ⬜ |
| Confidence recalibrations | S:3 / M:8 / L:15 | 0 | ⬜ |<HARD-GATE> Cannot exit until 80% of budget met. Print state ledger before each iteration decision. </HARD-GATE>
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | counterfactual-reasoning | Tactic for reasoning about what would happen if variables were different — supports causal identification and intervention analysis. | | evidence-weighing | Tactic for assessing the strength and relevance of evidence for causal claims — distinguishes correlation from causation. | | feedback-loop-detection | Tactic for identifying circular causation — detect feedback loops, classify as reinforcing or balancing, document loop structure. |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | model-gap-detection | SOP for finding gaps in the causal model — missing variables, unexplained effects, weak links. | | validation-report | SOP for generating a causal model validation report — summarize coverage, confidence, gaps, contradictions. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-19 | fail→pass | 24,236 | 32,162 | +33% | 1 | 1 | 0% | 2,935 | 5,392 | +84% | 0 | 0 | — |
case-06 | pass→fail | 34,064 | 68,782 | +102% | 1 | 1 | 0% | 2,394 | 8,917 | +272% | 0 | 0 | — |
case-01 | fail→pass | 17,223 | 43,857 | +155% | 1 | 1 | 0% | 1,187 | 7,989 | +573% | 0 | 0 | — |
case-02 | fail→pass | 40,237 | 41,478 | +3% | 1 | 1 | 0% | 6,430 | 8,107 | +26% | 0 | 0 | — |
case-03 | fail→pass | 46,558 | 48,381 | +4% | 1 | 1 | 0% | 6,734 | 8,955 | +33% | 0 | 0 | — |
case-04 | pass→fail | 36,877 | 45,233 | +23% | 1 | 1 | 0% | 3,504 | 7,057 | +101% | 0 | 0 | — |
case-05 | pass→fail | 21,303 | 47,293 | +122% | 1 | 1 | 0% | 2,824 | 8,933 | +216% | 0 | 0 | — |
case-07 | fail→pass | 49,035 | 56,097 | +14% | 1 | 1 | 0% | 5,096 | 7,560 | +48% | 0 | 0 | — |
case-08 | fail→pass | 39,170 | 37,953 | -3% | 1 | 1 | 0% | 3,078 | 4,912 | +60% | 0 | 0 | — |
case-09 | fail→pass | 40,488 | 60,038 | +48% | 1 | 1 | 0% | 2,926 | 8,935 | +205% | 0 | 0 | — |
case-10 | fail→pass | 42,501 | 57,552 | +35% | 1 | 1 | 0% | 2,556 | 7,477 | +193% | 0 | 0 | — |
case-11 | fail→pass | 42,822 | 45,661 | +7% | 1 | 1 | 0% | 6,867 | 8,362 | +22% | 0 | 0 | — |
case-12 | pass→pass | 36,018 | 48,878 | +36% | 1 | 1 | 0% | 4,352 | 8,481 | +95% | 0 | 0 | — |
case-13 | fail→pass | 32,546 | 51,923 | +60% | 1 | 1 | 0% | 4,270 | 8,922 | +109% | 0 | 0 | — |
case-14 | fail→pass | 15,984 | 24,459 | +53% | 1 | 1 | 0% | 3,269 | 4,424 | +35% | 0 | 0 | — |
case-15 | fail→pass | 23,695 | 44,812 | +89% | 1 | 1 | 0% | 3,100 | 8,219 | +165% | 0 | 0 | — |
case-16 | fail→pass | 31,654 | 35,602 | +12% | 1 | 1 | 0% | 4,566 | 7,381 | +62% | 0 | 0 | — |
case-17 | fail→pass | 16,285 | 24,616 | +51% | 1 | 1 | 0% | 1,737 | 4,970 | +186% | 0 | 0 | — |
case-18 | pass→pass | 30,604 | 50,485 | +65% | 1 | 1 | 0% | 4,557 | 8,913 | +96% | 0 | 0 | — |
case-20 | fail→pass | 25,762 | 45,157 | +75% | 1 | 1 | 0% | 3,913 | 8,508 | +117% | 0 | 0 | — |
case-21 | pass→pass | 24,289 | 35,787 | +47% | 1 | 1 | 0% | 3,124 | 6,917 | +121% | 0 | 0 | — |
case-22 | fail→pass | 34,171 | 39,143 | +15% | 1 | 1 | 0% | 4,891 | 8,912 | +82% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +59 percentage points is the difference between those two pass rates over the 22 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.