Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Strategy: Run BEFORE building any validator (sandbox/simulation/benchmark). Builds a non-circularity matrix of theory-claim × validator-assumption to detect when a validator would 'confirm' a theory only because it was built on the theory's own premises. A circular validator's PASS carries zero evidential weight. Methods: Cartwright nomological machines, Winsberg sanctioning-of-simulations, tautology detection.
.claude/skills/yogsoth-ai-circular-validation-audit/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 133% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 60% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 54% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 96% | 0% |
A pre-build gate for any validator we are about to construct — a sandbox, simulation, or benchmark that is supposed to verify a theory. The catastrophic failure mode: we build the validator using the theory's own assumptions as the validator's ground-truth machinery, run the theory's method on the validator's output, and celebrate when it "recovers" the truth — but the recovery was guaranteed by construction, not by the theory being right. The PASS carries zero evidential weight. This is the single most expensive mistake available, because a circular validator can burn enormous compute producing confident, meaningless confirmations. So this audit runs FIRST, on the validator DESIGN, before a line of it is built.
A validator V tests a theory T by: generating data D from a ground-truth generator G, running T's recovery method R on D, and checking if R(D) matches G's hidden truth. The test is non-circular only if G is NOT itself an instance of T's central assumptions. If G is built from the same functional factorization / the same noise law / the same mechanism that T claims to discover, then "R recovers G" is a tautology: we encoded the answer into the question. The audit's job is to find every place where G secretly contains T.
Rows = the theory's load-bearing claims/assumptions (the ones the validator is meant to test). Columns = the validator's ground-truth construction choices (how each mechanism is implemented, what functional forms or noise laws the generator uses, what intervention semantics the validator actually applies).
Each cell asks: does this construction choice ASSUME this claim?
| | validator: construction choice 1 | validator: construction choice 2 | validator: construction choice 3 | … | |---|---|---|---|---| | claim: load-bearing claim A | ? | ? | ? | ? | | claim: load-bearing claim B | ? | ? | ? | ? | | claim: load-bearing claim C | ? | ? | ? | ? | | … | ? | ? | ? | ? |
Fill rows with each load-bearing claim the validator must test. Fill columns with each ground-truth construction choice in the validator. Cell verdicts:
This is where "compute-first, non-circular" bites. Some claims may be provably unrecoverable from observational data alone — for those, a validator that "recovers" them must be using a genuine intervention. The audit must verify the intervention is real (the validator actually applies a do-intervention and the recovery uses interventional data) and not a relabeled observational shortcut. An intervention that's free in the validator is legitimate; an observational method wearing an interventional costume is circular.
NonCircularityMatrix: the filled matrix + per-red-cell de-circularization move (or "untestable here") + overall GREEN/REVISE verdict on the validator design + the explicit list of "truths the validator should make the theory FAIL on" (the adversarial ground-truth set that turns the validator from a confirmation machine into a real test).
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-07 | fail→pass | 8,600 | 12,412 | +44% | 1 | 1 | 0% | 1,480 | 3,441 | +133% | 0 | 0 | — |
case-02 | fail→fail | 27,933 | 28,067 | +0% | 1 | 1 | 0% | 4,630 | 5,723 | +24% | 0 | 0 | — |
case-03 | fail→pass | 28,897 | 23,621 | -18% | 1 | 1 | 0% | 4,777 | 5,255 | +10% | 0 | 0 | — |
case-04 | pass→pass | 11,226 | 32,273 | +187% | 1 | 1 | 0% | 2,048 | 6,796 | +232% | 0 | 0 | — |
case-05 | pass→pass | 18,365 | 25,320 | +38% | 1 | 1 | 0% | 2,935 | 5,267 | +79% | 0 | 0 | — |
case-06 | pass→pass | 21,309 | 24,634 | +16% | 1 | 1 | 0% | 4,991 | 6,724 | +35% | 0 | 0 | — |
case-01 | fail→fail | 28,761 | 30,215 | +5% | 1 | 1 | 0% | 4,991 | 6,785 | +36% | 0 | 0 | — |
case-08 | fail→pass | 12,969 | 13,652 | +5% | 1 | 1 | 0% | 2,136 | 3,420 | +60% | 0 | 0 | — |
case-09 | fail→pass | 15,111 | 14,971 | -1% | 1 | 1 | 0% | 2,315 | 3,570 | +54% | 0 | 0 | — |
case-10 | pass→fail | 14,927 | 17,888 | +20% | 1 | 1 | 0% | 2,414 | 4,161 | +72% | 0 | 0 | — |
case-11 | pass→pass | 15,330 | 16,923 | +10% | 1 | 1 | 0% | 2,229 | 3,943 | +77% | 0 | 0 | — |
case-12 | fail→fail | 33,714 | 11,588 | -66% | 1 | 1 | 0% | 2,218 | 3,121 | +41% | 0 | 0 | — |
case-13 | fail→pass | 16,230 | 23,425 | +44% | 1 | 1 | 0% | 2,596 | 5,077 | +96% | 0 | 0 | — |
case-14 | fail→pass | 13,166 | 8,661 | -34% | 1 | 1 | 0% | 1,944 | 2,550 | +31% | 0 | 0 | — |
case-15 | fail→pass | 8,633 | 2,735 | -68% | 1 | 1 | 0% | 1,351 | 1,653 | +22% | 0 | 0 | — |
case-16 | fail→pass | 12,569 | 7,356 | -41% | 1 | 1 | 0% | 1,999 | 2,443 | +22% | 0 | 0 | — |
case-17 | fail→pass | 10,145 | 4,297 | -58% | 1 | 1 | 0% | 1,600 | 1,962 | +23% | 0 | 0 | — |
case-18 | fail→fail | 15,552 | 18,617 | +20% | 1 | 1 | 0% | 2,524 | 4,543 | +80% | 0 | 0 | — |
case-19 | fail→pass | 24,157 | 4,554 | -81% | 1 | 1 | 0% | 1,928 | 2,037 | +6% | 0 | 0 | — |
case-20 | fail→pass | 10,118 | 2,539 | -75% | 1 | 1 | 0% | 1,527 | 1,687 | +10% | 0 | 0 | — |
case-21 | fail→pass | 15,041 | 12,134 | -19% | 1 | 1 | 0% | 2,195 | 3,112 | +42% | 0 | 0 | — |
case-22 | fail→fail | 6,424 | 1,828 | -72% | 1 | 1 | 0% | 1,065 | 1,535 | +44% | 0 | 0 | — |
case-23 | fail→pass | 15,705 | 1,702 | -89% | 1 | 1 | 0% | 2,586 | 1,545 | -40% | 0 | 0 | — |
case-24 | pass→pass | 18,433 | 18,082 | -2% | 1 | 1 | 0% | 2,843 | 4,005 | +41% | 0 | 0 | — |
case-25 | pass→pass | 16,437 | 14,899 | -9% | 1 | 1 | 0% | 2,717 | 3,717 | +37% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 23 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +48 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.